Stop tokens and stopping criteria
A model will keep writing forever unless something stops it, and the three mechanisms that stop it fail in different, specific, avoidable ways.
- 13 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A model does not know when it is finished. Something outside it has to decide, and there are three ways to decide.
Think about a tap. Water does not stop because the bucket is full. Water stops when somebody turns the tap, or the bucket overflows, or the supply is cut.
A language model is the tap. Left alone it produces one more token, then one more, without limit. Every ending you have ever seen was arranged.
The three ways to stop
The model asks to stop. During training it learned a special invisible token meaning "this turn is over". When it produces that token, generation ends. This is the clean ending, and it is the one you want.
You set a limit. A maximum number of tokens. When the count is reached, generation ends wherever it happens to be, possibly mid-sentence. This is a safety net, not a plan.
You watch for a phrase. You tell the server to stop as soon as certain text appears. Say, the word that marks the start of the next speaker's turn. This is a stop string.
Every ending in every chat product you have used is one of these three.
Why the third one is harder than it looks
Models do not write letters. They write chunks of text of varying size.
The phrase you want to stop on might be spread across three chunks. No single chunk equals your phrase, so a check that compares each chunk to your phrase never fires. The model runs to the limit and your stop string is ignored.
There is a second, subtler version of the same trouble. Even if you catch the phrase correctly, a streaming interface may already have shown it. The user saw the text before you noticed. The user sees the text you were trying to hide.
The fix for both is to hold back the last few characters until they are proven safe. Not difficult, and frequently missing.
How it looks
chunks arriving: "Sunday" "." "\n" "User" ":"
compare each chunk to "\nUser:" -> never matches. Never stops.
watch the running text
"...Sunday." safe
"...Sunday.\n" might be the start of the phrase. hold it back.
"...Sunday.\nUser" still might be. hold it back.
"...Sunday.\nUser:" it is. stop, and do not show the held-back part.Why your model rambles
Two very common causes, both fixable.
The wrong template. Chat models expect their conversation formatted a particular way, with particular invisible markers. Format it wrongly and the model never reaches the state where it would produce its ending token. It keeps going, often inventing the next turn of the conversation by itself.
The wrong ending token. Some models have two: one meaning "the document is over", another meaning "my turn is over". A tool configured with only the first will run past every turn boundary. This produces the classic bug where a model answers your question and then invents your next question.
Where you have already seen this
- A chat reply cut off mid-word, with a "continue" button.
- A local model that answers and then role-plays the whole conversation alone.
- An API response with a field saying whether it finished naturally or hit the limit.
- A model that ends every answer with a phrase you never asked for.
Remember this
- Nothing stops on its own: either the model asks, or a limit hits, or you watch for a phrase.
- A stop phrase can be split across chunks, so a per-chunk comparison never fires.
- Rambling is usually the wrong chat template or the wrong ending token, not a bad model.
What to learn next
- Why temperature zero is still not deterministic — the other surprise in a supposedly fixed pipeline.
- Streaming tokens correctly — the hold-back buffer in its natural home.
- Tokenization — why your stop string does not line up with token boundaries.
Developer — Code and libraries.
Setup
# no libraries needed
python3 --versionThree stop-string implementations against the same token stream. Two of them are what people actually ship.
Three ways to handle a stop string, two of them wrong
# What a tokenizer actually hands you, one piece at a time.
STREAM = ["Sure", ",", " the", " office", " is", " closed", " on", " Sunday", ".",
"\n", "User", ":", " what", " about", " Monday", "?"]
STOP = "\nUser:"
MAX_NEW = 32
def naive_token_match():
"""Compare each token to the stop string. Popular, and wrong."""
out = ""
for t in STREAM[:MAX_NEW]:
if t == STOP:
return out, "stop"
out += t
return out, "length"
def endswith_after_append():
"""Correct stopping point, but the stop text was already emitted."""
out = ""
emitted = ""
for t in STREAM[:MAX_NEW]:
out += t
emitted += t # this is what a streaming client saw
if out.endswith(STOP):
return emitted, "stop"
return emitted, "length"
def held_back():
"""Hold back the last len(STOP)-1 characters until they are proven safe."""
buf, emitted = "", ""
for t in STREAM[:MAX_NEW]:
buf += t
if STOP in buf:
emitted += buf[: buf.index(STOP)]
return emitted, "stop"
safe = len(buf) - (len(STOP) - 1) # anything older than this cannot start STOP
if safe > 0:
emitted += buf[:safe]
buf = buf[safe:]
return emitted + buf, "length"
for name, fn in (("token equality ", naive_token_match),
("endswith, no trim", endswith_after_append),
("held-back buffer ", held_back)):
text, reason = fn()
print(f"{name}: finish_reason={reason:<7} user saw {text!r}")
print("\nwhy token equality can never work here:")
running = ""
for t in STREAM[:12]:
running += t
flag = " <- the stop string is now complete" if running.endswith(STOP) else ""
print(f" token {t!r:<10} -> text so far ends {running[-8:]!r}{flag}")
print(f"\nThe stop string {STOP!r} is spread over the tokens "
f"{STREAM[9]!r}, {STREAM[10]!r} and {STREAM[11]!r}.")
print("No single token equals it, so a per-token comparison never fires.")token equality : finish_reason=length user saw 'Sure, the office is closed on Sunday.\nUser: what about Monday?' endswith, no trim: finish_reason=stop user saw 'Sure, the office is closed on Sunday.\nUser:' held-back buffer : finish_reason=stop user saw 'Sure, the office is closed on Sunday.' why token equality can never work here: token 'Sure' -> text so far ends 'Sure' token ',' -> text so far ends 'Sure,' token ' the' -> text so far ends 'ure, the' token ' office' -> text so far ends 'e office' token ' is' -> text so far ends 'ffice is' token ' closed' -> text so far ends 's closed' token ' on' -> text so far ends 'losed on' token ' Sunday' -> text so far ends 'n Sunday' token '.' -> text so far ends ' Sunday.' token '\n' -> text so far ends 'Sunday.\n' token 'User' -> text so far ends 'ay.\nUser' token ':' -> text so far ends 'y.\nUser:' <- the stop string is now complete The stop string '\nUser:' is spread over the tokens '\n', 'User' and ':'. No single token equals it, so a per-token comparison never fires.
Reading the output
Token equality produced the whole hallucinated next turn. finish_reason says length, so the caller believes the model was cut off and offers a "continue" button. Nothing was cut off. The stop string was there and the check could not see it. This is the single most common stop-sequence bug.
endswith stopped correctly and leaked \nUser: to the user. In a non-streaming API you can strip it afterwards, and most servers do. In a streaming API those characters have already gone down the wire. Trimming after the fact is not possible once the bytes have left.
The held-back version got both right. Text ends at Sunday. and nothing extra reached the client. The rule is short. Any character older than len(STOP) - 1 from the end of the buffer cannot start the stop string, so it is safe to emit.
The per-token walk shows the mechanism. The stop string spans three tokens: '\n', 'User' and ':'. A tokenizer never guarantees your phrase aligns with token boundaries, and there is no setting that makes it. The check must be over characters.
The three stopping mechanisms, and their real names
out = model.generate(
**inputs,
max_new_tokens=256, # the hard limit
eos_token_id=[128001, 128009], # a LIST: end-of-text and end-of-turn
stop_strings=["\nUser:", "</answer>"],
tokenizer=tok, # required when stop_strings is used
)Three notes on that call.
eos_token_id accepts a list, and for chat models it usually needs one. Llama 3 Instruct has <|end_of_text|> from pretraining and <|eot_id|> marking the end of a turn. A tool configured with only the first runs straight past every turn boundary and starts writing the user's next message.
stop_strings requires you to pass tokenizer=, because matching happens over decoded text rather than ids. Omitting it raises an error, which is a good design.
max_new_tokens counts generated tokens only. max_length counts prompt plus generation, and mixing them up is why a long prompt sometimes produces a one-token answer.
finish_reason, and why it is the field to log
OpenAI-compatible servers return one of:
| Value | Meaning | What to do |
|---|---|---|
stop | Hit an EOS token or a stop string | Normal |
length | Hit max_tokens | The answer is truncated. Do not parse it as complete |
tool_calls | Stopped to call a tool | Dispatch the call |
content_filter | Blocked by a safety filter | Surface it, do not retry blindly |
Truncated JSON is the classic downstream failure: finish_reason says length, the code ignores it, the parser throws, and the retry produces the same truncation. Check this field before parsing anything.
Common mistakes
Hand-building the chat prompt. Write "User: ...\nAssistant: " yourself and the model never sees the special tokens it was trained to end on. Use tokenizer.apply_chat_template(messages, add_generation_prompt=True) and the ending token arrives on its own.
A stop string that occurs in legitimate output. Stopping on "\n\n" truncates every answer with a paragraph break. Stopping on "``"` truncates every code block at its opening fence.
Forgetting that stop strings are stripped, inconsistently. OpenAI excludes the stop sequence from the output. Some servers include it. Do not build a parser that depends on either behaviour.
No max_tokens at all. A model in a repetition loop with an unbounded limit costs real money and holds a KV cache slot until it hits the context limit. Always set a ceiling.
Stopping on a token id instead of text. Token ids are model-specific and change between checkpoints and tokenizer versions. Hard-coding 50256 works until the day you swap models.
Try it yourself
Change STOP to "Sunday" and re-run. All three implementations now stop, because "Sunday" happens to be exactly one token. That is precisely why per-token matching seems to work in testing and fails in production. Then set STOP = "day." and watch only the held-back version behave.
What to learn next
- Why temperature zero is still not deterministic — the other surprise in a supposedly fixed pipeline.
- Streaming tokens correctly — the hold-back buffer in its natural home.
- Tokenization — why your stop string does not line up with token boundaries.
Researcher — Mathematics and papers.
The three mechanisms, precisely
Generation halts when any element of a stopping-criteria set returns true after a step:
$$ \text{halt}(y_{1:t}) = \bigvee_i C_i(y_{1:t}) $$
In transformers these are EosTokenCriteria, MaxLengthCriteria, MaxTimeCriteria, StopStringCriteria and ConfidenceCriteria, composed in a StoppingCriteriaList. Under batching, criteria are evaluated per sequence, and finished rows are padded until the whole batch completes. That is why a batched call can be slower than the sum of its parts, and why continuous batching exists.
EOS is a training artefact, not a property of language
The end-of-sequence token is learned. Its probability at position $t$ is the model's estimate of the chance that a well-formed document ends there, under the distribution of documents it was trained on.
Three consequences worth stating.
Format mismatch destroys it. If the inference-time prompt format differs from the training format, the model is off-distribution and $P(\text{EOS})$ is unreliable. Rambling is the visible symptom of an off-distribution prompt, and it is a formatting bug, not a decoding bug.
Truncation biases it. Pretraining corpora chunked to a fixed context length teach the model that documents rarely end. Length-control methods and instruction tuning are partly corrections for this.
Sampling can suppress it. EOS is one token competing with the whole vocabulary. Under high temperature its probability shrinks along with everything else's, so hot sampling produces measurably longer outputs. min_new_tokens works by forcing $P(\text{EOS}) = 0$ for the first $n$ steps, which is the same mechanism used in reverse.
The multiple-EOS problem
Modern chat models separate document-level and turn-level termination:
| Model family | Document end | Turn end |
|---|---|---|
| Llama 3 | `< | end_of_text |
| Qwen 2.5 / 3 | `< | endoftext |
| Gemma | <eos> | <end_of_turn> |
The correct configuration lists both. A model card's generation_config.json is the authority, and a serving stack that reads only config.eos_token_id will miss the turn-level token. This single misconfiguration accounts for a large share of "the model talks to itself" reports.
Stop-string matching over tokens
Matching decoded text is correct but naive implementations decode the whole sequence every step, which is $O(n^2)$ over a generation.
transformers implements StopStringCriteria with a precomputed structure instead. For each stop string it records, per vocabulary token, the valid overlaps between that token's text and suffixes of the stop string. Matching then becomes a tensor operation over the last few tokens, rather than a decode-and-search. The cost is a one-time build against the vocabulary.
The streaming discipline is separate from matching and is the part usually missing. A correct streaming server holds back $\max_i(|s_i|) - 1$ characters, where $s_i$ ranges over the stop strings, before emitting. Held-back text is flushed on EOS or on the limit. Without this, a stop sequence is visible to the client before the server has decided to stop.
Interaction with other decoding machinery
Speculative decoding. A draft block may contain the stop token in its middle. The verifier must truncate at that position and discard the rest, rather than accepting the whole accepted prefix. Implementations that check for EOS only after applying a full block emit tokens past the stop.
Constrained decoding. A grammar that never reaches a terminal state makes EOS illegal forever, so generation runs to max_tokens. Every grammar needs a reachable accepting state, and this is the most common grammar bug.
Structured output. With a JSON grammar the closing brace is the true terminator and EOS may be irrelevant. Stop strings and grammars are alternative termination mechanisms; using both without care produces truncated JSON.
Tool calling. The model stops at a tool-call boundary and the response carries finish_reason: tool_calls, not stop. Agent loops that branch only on stop silently drop tool calls.
Billing and capacity
max_tokens is a cost control and a capacity control at once. It bounds the per-request bill, bounds the KV cache a request can hold, and lets a scheduler reason about worst-case memory. Under paged allocation the reservation is on demand, so a generous max_tokens is cheaper than it used to be. It still bounds preemption risk, and it still bounds the invoice.
Further reading
- HuggingFace
transformers—generation/stopping_criteria.py, the reference implementation of stop-string matching. - The
generation_config.jsonof any chat model, which is where the real EOS configuration lives. - OpenAI API reference —
finish_reasonsemantics, which every compatible server mirrors.
What to learn next
- Why temperature zero is still not deterministic — the other surprise in a supposedly fixed pipeline.
- Streaming tokens correctly — the hold-back buffer in its natural home.
- Tokenization — why your stop string does not line up with token boundaries.