Streaming tokens correctly
Sending tokens as they appear is easy to get almost right, and the two ways it usually goes wrong both hit non-English text and both are fixed by holding bytes back.
- 13 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Streaming means sending each piece of the answer as soon as it exists, instead of waiting for the whole thing.
Think about a meal at home. Nobody waits until every dish is finished before anyone eats. The rotis come to the table hot, one at a time, while the rest is still cooking.
Nothing got faster. The waiting became bearable, because something arrived immediately.
That is exactly what streaming does for a model. The full answer takes the same total time. You start reading in half a second instead of staring at a blank screen for twelve.
What is actually being sent
The model produces small chunks of text, one at a time. The server sends each one down an open connection as it appears, and your screen appends it.
That is the whole design. One connection, held open, with small updates flowing along it, until a final message says the answer is complete.
The problem nobody expects
Those chunks are not letters and not words. They are pieces of the underlying byte-level storage of text.
For English that mostly does not matter. A chunk is usually a word or part of a word, and showing it straight away works.
For Hindi, Marathi, Tamil, Bengali, Chinese, or an emoji, a single character takes several bytes to store. A chunk boundary can land in the middle of one character.
Show that half-character on its own and you get the black diamond with a question mark, or a crash. Users who write in Devanagari see this constantly on badly built tools. Users who write in English never see it, and do not believe it happens.
How it looks
the word: चाय
stored as bytes: [.. .. ..][.. .. ..][.. .. ..]
chunk boundary lands here: ^
show each chunk immediately -> error, or a broken square
hold the incomplete part back -> correct text, one chunk laterThe fix is one sentence. If the bytes you are holding do not form a complete character, wait. Show nothing until the next chunk arrives.
The second problem, which is quieter
There is a related trap in how chunks become readable text.
Some models store a space as part of the following chunk. That space vanishes if you convert each chunk on its own. Convert them individually and join the results, and every space disappears.
"The chai is ready" becomes "Thechaiisready". The correct method is to convert the whole answer so far and send only the part that is new.
Where you have already seen this
- ChatGPT typing out an answer word by word.
- A voice assistant speaking a sentence before it has planned the next.
- A translation tool filling in a long paragraph progressively.
- A badly built chat tool showing black diamonds in Hindi, then fixing itself a moment later.
Remember this
- Streaming does not make the answer faster; it makes the first word arrive sooner.
- A chunk boundary can split a character in Indian languages, and in emoji.
- Hold back incomplete bytes, and build text from the whole answer so far rather than chunk by chunk.
What to learn next
- Stop tokens and stopping criteria — the other reason a streaming server holds bytes back.
- Tokenization — why token boundaries and character boundaries disagree.
- Model deployment — putting the whole pipeline behind a real endpoint.
Developer — Code and libraries.
Setup
# no libraries needed for the first block
python3 --versionThe whole problem is byte boundaries. Python's standard library already has the correct tool, and most streaming bugs come from not knowing it exists.
Incremental decoding, and what happens without it
import codecs, json, sys, io
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding="utf-8")
TEXT = "आपका चाय ☕ तैयार है"
raw = TEXT.encode("utf-8")
# Byte-level tokenizers cut on byte boundaries, not on character boundaries.
CHUNKS = [raw[i:i + 5] for i in range(0, len(raw), 5)]
print(f"text : {TEXT}")
print(f"characters : {len(TEXT)}")
print(f"utf-8 bytes : {len(raw)}")
print(f"pretend tokens : {len(CHUNKS)} pieces of at most 5 bytes\n")
print("decoding each piece on its own:")
for i, c in enumerate(CHUNKS):
try:
print(f" piece {i}: {c.decode('utf-8')!r}")
except UnicodeDecodeError as e:
print(f" piece {i}: UnicodeDecodeError -> {e.reason}")
print("\nsame pieces through an incremental decoder:")
dec = codecs.getincrementaldecoder("utf-8")()
out = ""
for i, c in enumerate(CHUNKS):
piece = dec.decode(c) # holds back an unfinished character
out += piece
print(f" piece {i}: emitted {piece!r}")
print(f" final: {out!r}")
print(f" matches the original: {out == TEXT}")
print("\nthe same bytes as server-sent events, which is what the wire looks like:")
dec = codecs.getincrementaldecoder("utf-8")()
for c in CHUNKS[:4]:
piece = dec.decode(c)
if not piece:
continue # nothing complete yet: send NOTHING
frame = {"choices": [{"delta": {"content": piece}, "finish_reason": None}]}
print("data: " + json.dumps(frame, ensure_ascii=False))
print('data: {"choices": [{"delta": {}, "finish_reason": "stop"}]}')
print("data: [DONE]")text : आपका चाय ☕ तैयार है
characters : 19
utf-8 bytes : 49
pretend tokens : 10 pieces of at most 5 bytes
decoding each piece on its own:
piece 0: UnicodeDecodeError -> unexpected end of data
piece 1: UnicodeDecodeError -> invalid start byte
piece 2: UnicodeDecodeError -> invalid start byte
piece 3: UnicodeDecodeError -> invalid start byte
piece 4: UnicodeDecodeError -> invalid start byte
piece 5: UnicodeDecodeError -> invalid start byte
piece 6: UnicodeDecodeError -> unexpected end of data
piece 7: UnicodeDecodeError -> invalid start byte
piece 8: UnicodeDecodeError -> invalid start byte
piece 9: UnicodeDecodeError -> invalid start byte
same pieces through an incremental decoder:
piece 0: emitted 'आ'
piece 1: emitted 'पक'
piece 2: emitted 'ा '
piece 3: emitted 'चा'
piece 4: emitted 'य '
piece 5: emitted '☕ त'
piece 6: emitted 'ै'
piece 7: emitted 'या'
piece 8: emitted 'र '
piece 9: emitted 'है'
final: 'आपका चाय ☕ तैयार है'
matches the original: True
the same bytes as server-sent events, which is what the wire looks like:
data: {"choices": [{"delta": {"content": "आ"}, "finish_reason": null}]}
data: {"choices": [{"delta": {"content": "पक"}, "finish_reason": null}]}
data: {"choices": [{"delta": {"content": "ा "}, "finish_reason": null}]}
data: {"choices": [{"delta": {"content": "चा"}, "finish_reason": null}]}
data: {"choices": [{"delta": {}, "finish_reason": "stop"}]}
data: [DONE]Reading the output
Nineteen characters take forty-nine bytes. Each Devanagari character costs three. The same sentence in English would be close to one byte per character. Every byte-level budget written by an English-speaking developer is wrong for Indian languages by roughly a factor of three, and that includes token budgets.
All ten pieces failed to decode on their own. Not one, not a couple — every single one. That is what a naive chunk.decode("utf-8") in a streaming loop does to Hindi.
The incremental decoder emitted nothing extra and everything correct. piece 0 gave one character out of five bytes and silently kept the remaining two. piece 5 gave three characters, including one that began in an earlier chunk. The reconstruction matches the original exactly.
Piece 6 emitted a combining mark on its own. The ै is a vowel sign that attaches to the previous consonant. It is a valid complete character, so the decoder correctly emits it, and a renderer will attach it to what came before. Byte-level correctness is not the same as visual stability, and a client that animates each delta separately can make text appear to shuffle. Byte correctness is the part you must get right; visual smoothness is a client-side choice.
The if not piece: continue line is the streaming rule. When the decoder holds bytes back it returns an empty string. Sending an empty delta frame is not harmful, but sending a broken one is, and the guard makes the intent explicit.
The second bug: building text from tokens
import sys, io
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding="utf-8")
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("TinyLlama/TinyLlama-1.1B-Chat-v1.0")
ids = tok("The chai is ready", add_special_tokens=False).input_ids
print("token ids:", ids)
print("per-token decode :", [tok.decode([i]) for i in ids])
print("joined :", repr("".join(tok.decode([i]) for i in ids)))
print("whole-sequence :", repr(tok.decode(ids)))
out, deltas = "", []
for n in range(1, len(ids) + 1):
full = tok.decode(ids[:n]) # decode the whole prefix
deltas.append(full[len(out):]) # send only what is new
out = full
print("prefix-diff deltas:", deltas)
print("joined deltas :", repr("".join(deltas)))token ids: [450, 521, 1794, 338, 7960] per-token decode : ['The', 'ch', 'ai', 'is', 'ready'] joined : 'Thechaiisready' whole-sequence : 'The chai is ready' prefix-diff deltas: ['The', ' ch', 'ai', ' is', ' ready'] joined deltas : 'The chai is ready'
This downloads the tokenizer files only, around a megabyte, not the model weights. Run against transformers 5.6.2.
'Thechaiisready'. Every space is gone. This tokenizer marks a leading space with a special symbol inside the token, and decoding one token at a time strips it as leading whitespace. Decoding tokens individually and concatenating is wrong for every SentencePiece-derived tokenizer, which is most of the Llama, Mistral and Gemma families.
The prefix-diff deltas are correct. 'The', ' ch', 'ai', ' is', ' ready' — spaces attached to the right tokens, and joining them reproduces the sentence. Decoding the full prefix each step is $O(n^2)$ over a long generation, so production detokenizers keep a small window of recent tokens and diff within it. The principle is the same: decode a span, not a single token.
What a real streaming server holds back
Three separate buffers, and each exists for its own reason:
- Incomplete UTF-8 bytes — until the character is whole.
- Possible stop-string prefixes — up to
len(stop) - 1characters, covered in stopping criteria. - Tokenizer context — a few recent tokens so the decoder knows how to join.
All three are hold-backs. All three are missing from most hand-rolled streaming loops.
The wire format
Server-sent events, which is HTTP with a specific body shape:
Content-Type: text/event-stream
Cache-Control: no-cache
X-Accel-Buffering: no
data: {"choices":[{"delta":{"content":"The"}}]}
data: {"choices":[{"delta":{"content":" chai"}}]}
data: [DONE]Each frame is data: plus one JSON object, terminated by a blank line. [DONE] is a convention rather than part of the SSE specification.
X-Accel-Buffering: no matters more than it looks. Nginx buffers proxied responses by default. A correct streaming server behind a default reverse proxy then delivers the whole answer in one lump. This is the most common "streaming does not work in production but works locally" cause.
Common mistakes
Decoding each token alone. Shown above. Spaces vanish, and non-English text breaks worse.
No client-side cancellation. If the user closes the tab, the server should stop generating. Without this you pay for tokens nobody will read and you hold a KV cache slot.
Buffering in the framework. Flask, Django and Express all need explicit streaming responses. A generator that is silently collected into a string is easy to write by accident.
Streaming a response that must be parsed as a whole. For JSON output, streaming gives the user a growing blob of syntax. Stream it only if you have a partial-JSON parser on the client, and check finish_reason before parsing.
Measuring only total time. Time to first token and inter-token latency are what streaming changes. Total time is barely affected. Log all three separately.
Try it yourself
Change CHUNKS to slice at 1 byte instead of 5. Every piece now fails the naive decode and the incremental decoder still reconstructs the sentence exactly. Then try a string of pure ASCII and watch the naive version work perfectly — which is precisely how this bug survives testing and ships.
What to learn next
- Stop tokens and stopping criteria — the other reason a streaming server holds bytes back.
- Tokenization — why token boundaries and character boundaries disagree.
- Model deployment — putting the whole pipeline behind a real endpoint.
Researcher — Mathematics and papers.
Why detokenization is not the inverse of tokenization
Tokenizers are not injective at the piece level. Three representative behaviours:
- Byte-level BPE (GPT-2, Llama 3, Qwen) maps bytes into a printable alias alphabet, so a token's string may hold half a UTF-8 sequence. Only concatenation and a final byte-level decode reconstruct the text.
- SentencePiece unigram and BPE (Llama 2, Mistral, Gemma) encode a leading space as
▁. Decoding a single token strips it as leading whitespace, so per-token decode loses word boundaries — the failure reproduced above. - WordPiece (BERT) marks continuation with
##and requires the opposite join rule.
The general contract is therefore: detokenize a span, not a token. Streaming detokenizers keep a rolling window of recent token ids, decode the window, and emit the difference against what has already been sent. vLLM's incremental detokenizer keeps a small prefix_offset and read_offset pair over the token list. That avoids the $O(n^2)$ cost of decoding the whole prefix each step.
Grapheme clusters, which the byte level does not solve
Correct UTF-8 decoding gives valid code points. It does not give stable user-perceived characters.
A Devanagari syllable such as त plus ै is two code points and one grapheme cluster. Emoji sequences are worse: a family emoji is several code points joined by zero-width joiners, and a flag is two regional-indicator code points.
Splitting a delta between code points of one cluster is byte-valid and visually unstable. The client renders an intermediate form, which changes when the next delta arrives. Whether to buffer to grapheme-cluster boundaries is a client decision with a latency cost, and correctness at the byte level is the server's responsibility either way.
The Unicode text-segmentation algorithm (UAX #29) defines the cluster boundaries. Python's standard library does not implement it; regex with \X does.
Latency decomposition
$$ T_{\text{total}} = T_{\text{queue}} + T_{\text{prefill}} + \sum_{t=2}^{N} T_{\text{decode}, t} $$
Streaming reveals $T_{\text{queue}} + T_{\text{prefill}}$ as time to first token and $T_{\text{decode}}$ as inter-token latency. It does not reduce either. What it changes is perceived latency. The threshold that matters is human reading speed. Roughly 4 to 6 tokens per second is readable. Beyond about 20 tokens per second, extra speed is invisible to a reader. It is not invisible to an agent consuming the output.
Report the distribution, not the mean. A p99 TTFT of eight seconds with a p50 of four hundred milliseconds is a product problem that a mean hides completely.
Transport choices
| Transport | Direction | Reconnect | Notes |
|---|---|---|---|
| SSE over HTTP | server to client | built into the browser API | The de facto standard for LLM APIs |
| WebSocket | both ways | manual | Needed for interruption and voice |
| HTTP chunked transfer | server to client | none | What SSE is built on |
| gRPC streaming | both ways | built in | Common between internal services |
SSE won for LLM APIs because it is plain HTTP, survives proxies and CDNs, and needs no protocol upgrade. Its constraints are real: text only, one direction, and a browser limit on concurrent connections per origin over HTTP/1.1 — six, which HTTP/2 removes by multiplexing.
Interaction with the rest of the decoding stack
Speculative decoding produces several tokens in one step, so deltas arrive in bursts rather than evenly. Inter-token latency becomes bimodal, and a mean hides that. Some servers deliberately pace bursts to keep the visual rhythm even.
Stop strings require the hold-back described above. A server that streams before checking has already leaked the stop text.
Structured output streams syntactically valid prefixes that are not parseable JSON. Either buffer to completion or use a partial parser on the client.
Cancellation must propagate from a closed client connection to the scheduler, or the request keeps generating and keeps holding KV blocks. In vLLM this is an abort on the request id; without it, an abandoned long generation is pure waste.
Further reading
- WHATWG HTML specification, Server-sent events — the transport, including reconnection semantics.
- Unicode Standard Annex #29, Unicode Text Segmentation — grapheme cluster boundaries.
- vLLM source,
detokenize_incrementally— the reference incremental detokenizer. - HuggingFace
transformers,TextIteratorStreamer— the simplest correct streaming implementation to read.
What to learn next
- Stop tokens and stopping criteria — the other reason a streaming server holds bytes back.
- Tokenization — why token boundaries and character boundaries disagree.
- Model deployment — putting the whole pipeline behind a real endpoint.