How Text Is Generated

Streaming tokens correctly

Sending tokens as they appear is easy to get almost right, and the two ways it usually goes wrong both hit non-English text and both are fixed by holding bytes back.

Read these first

On this page 7
  1. What is actually being sent
  2. The problem nobody expects
  3. How it looks
  4. The second problem, which is quieter
  5. Where you have already seen this
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Streaming means sending each piece of the answer as soon as it exists, instead of waiting for the whole thing.

Think about a meal at home. Nobody waits until every dish is finished before anyone eats. The rotis come to the table hot, one at a time, while the rest is still cooking.

Nothing got faster. The waiting became bearable, because something arrived immediately.

That is exactly what streaming does for a model. The full answer takes the same total time. You start reading in half a second instead of staring at a blank screen for twelve.

What is actually being sent

The model produces small chunks of text, one at a time. The server sends each one down an open connection as it appears, and your screen appends it.

That is the whole design. One connection, held open, with small updates flowing along it, until a final message says the answer is complete.

The problem nobody expects

Those chunks are not letters and not words. They are pieces of the underlying byte-level storage of text.

For English that mostly does not matter. A chunk is usually a word or part of a word, and showing it straight away works.

For Hindi, Marathi, Tamil, Bengali, Chinese, or an emoji, a single character takes several bytes to store. A chunk boundary can land in the middle of one character.

Show that half-character on its own and you get the black diamond with a question mark, or a crash. Users who write in Devanagari see this constantly on badly built tools. Users who write in English never see it, and do not believe it happens.

How it looks

   the word: चाय

   stored as bytes:  [.. .. ..][.. .. ..][.. .. ..]
   chunk boundary lands here:      ^

   show each chunk immediately  ->  error, or a broken square
   hold the incomplete part back ->  correct text, one chunk later

The fix is one sentence. If the bytes you are holding do not form a complete character, wait. Show nothing until the next chunk arrives.

The second problem, which is quieter

There is a related trap in how chunks become readable text.

Some models store a space as part of the following chunk. That space vanishes if you convert each chunk on its own. Convert them individually and join the results, and every space disappears.

"The chai is ready" becomes "Thechaiisready". The correct method is to convert the whole answer so far and send only the part that is new.

Where you have already seen this

  • ChatGPT typing out an answer word by word.
  • A voice assistant speaking a sentence before it has planned the next.
  • A translation tool filling in a long paragraph progressively.
  • A badly built chat tool showing black diamonds in Hindi, then fixing itself a moment later.

Remember this

  • Streaming does not make the answer faster; it makes the first word arrive sooner.
  • A chunk boundary can split a character in Indian languages, and in emoji.
  • Hold back incomplete bytes, and build text from the whole answer so far rather than chunk by chunk.

What to learn next

Developer — Code and libraries.

Setup

bash
# no libraries needed for the first block
python3 --version

The whole problem is byte boundaries. Python's standard library already has the correct tool, and most streaming bugs come from not knowing it exists.

Incremental decoding, and what happens without it

streaming.py
import codecs, json, sys, io
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding="utf-8")

TEXT = "आपका चाय ☕ तैयार है"
raw = TEXT.encode("utf-8")
# Byte-level tokenizers cut on byte boundaries, not on character boundaries.
CHUNKS = [raw[i:i + 5] for i in range(0, len(raw), 5)]

print(f"text            : {TEXT}")
print(f"characters      : {len(TEXT)}")
print(f"utf-8 bytes     : {len(raw)}")
print(f"pretend tokens  : {len(CHUNKS)} pieces of at most 5 bytes\n")

print("decoding each piece on its own:")
for i, c in enumerate(CHUNKS):
    try:
        print(f"   piece {i}: {c.decode('utf-8')!r}")
    except UnicodeDecodeError as e:
        print(f"   piece {i}: UnicodeDecodeError -> {e.reason}")

print("\nsame pieces through an incremental decoder:")
dec = codecs.getincrementaldecoder("utf-8")()
out = ""
for i, c in enumerate(CHUNKS):
    piece = dec.decode(c)                 # holds back an unfinished character
    out += piece
    print(f"   piece {i}: emitted {piece!r}")
print(f"   final: {out!r}")
print(f"   matches the original: {out == TEXT}")

print("\nthe same bytes as server-sent events, which is what the wire looks like:")
dec = codecs.getincrementaldecoder("utf-8")()
for c in CHUNKS[:4]:
    piece = dec.decode(c)
    if not piece:
        continue                          # nothing complete yet: send NOTHING
    frame = {"choices": [{"delta": {"content": piece}, "finish_reason": None}]}
    print("data: " + json.dumps(frame, ensure_ascii=False))
print('data: {"choices": [{"delta": {}, "finish_reason": "stop"}]}')
print("data: [DONE]")
Output
text            : आपका चाय ☕ तैयार है
characters      : 19
utf-8 bytes     : 49
pretend tokens  : 10 pieces of at most 5 bytes

decoding each piece on its own:
   piece 0: UnicodeDecodeError -> unexpected end of data
   piece 1: UnicodeDecodeError -> invalid start byte
   piece 2: UnicodeDecodeError -> invalid start byte
   piece 3: UnicodeDecodeError -> invalid start byte
   piece 4: UnicodeDecodeError -> invalid start byte
   piece 5: UnicodeDecodeError -> invalid start byte
   piece 6: UnicodeDecodeError -> unexpected end of data
   piece 7: UnicodeDecodeError -> invalid start byte
   piece 8: UnicodeDecodeError -> invalid start byte
   piece 9: UnicodeDecodeError -> invalid start byte

same pieces through an incremental decoder:
   piece 0: emitted 'आ'
   piece 1: emitted 'पक'
   piece 2: emitted 'ा '
   piece 3: emitted 'चा'
   piece 4: emitted 'य '
   piece 5: emitted '☕ त'
   piece 6: emitted 'ै'
   piece 7: emitted 'या'
   piece 8: emitted 'र '
   piece 9: emitted 'है'
   final: 'आपका चाय ☕ तैयार है'
   matches the original: True

the same bytes as server-sent events, which is what the wire looks like:
data: {"choices": [{"delta": {"content": "आ"}, "finish_reason": null}]}
data: {"choices": [{"delta": {"content": "पक"}, "finish_reason": null}]}
data: {"choices": [{"delta": {"content": "ा "}, "finish_reason": null}]}
data: {"choices": [{"delta": {"content": "चा"}, "finish_reason": null}]}
data: {"choices": [{"delta": {}, "finish_reason": "stop"}]}
data: [DONE]

Reading the output

Nineteen characters take forty-nine bytes. Each Devanagari character costs three. The same sentence in English would be close to one byte per character. Every byte-level budget written by an English-speaking developer is wrong for Indian languages by roughly a factor of three, and that includes token budgets.

All ten pieces failed to decode on their own. Not one, not a couple — every single one. That is what a naive chunk.decode("utf-8") in a streaming loop does to Hindi.

The incremental decoder emitted nothing extra and everything correct. piece 0 gave one character out of five bytes and silently kept the remaining two. piece 5 gave three characters, including one that began in an earlier chunk. The reconstruction matches the original exactly.

Piece 6 emitted a combining mark on its own. The ै is a vowel sign that attaches to the previous consonant. It is a valid complete character, so the decoder correctly emits it, and a renderer will attach it to what came before. Byte-level correctness is not the same as visual stability, and a client that animates each delta separately can make text appear to shuffle. Byte correctness is the part you must get right; visual smoothness is a client-side choice.

The if not piece: continue line is the streaming rule. When the decoder holds bytes back it returns an empty string. Sending an empty delta frame is not harmful, but sending a broken one is, and the guard makes the intent explicit.

The second bug: building text from tokens

detokenize.py
import sys, io
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding="utf-8")
from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("TinyLlama/TinyLlama-1.1B-Chat-v1.0")
ids = tok("The chai is ready", add_special_tokens=False).input_ids

print("token ids:", ids)
print("per-token decode :", [tok.decode([i]) for i in ids])
print("joined           :", repr("".join(tok.decode([i]) for i in ids)))
print("whole-sequence   :", repr(tok.decode(ids)))

out, deltas = "", []
for n in range(1, len(ids) + 1):
    full = tok.decode(ids[:n])          # decode the whole prefix
    deltas.append(full[len(out):])      # send only what is new
    out = full
print("prefix-diff deltas:", deltas)
print("joined deltas    :", repr("".join(deltas)))
Output
token ids: [450, 521, 1794, 338, 7960]
per-token decode : ['The', 'ch', 'ai', 'is', 'ready']
joined           : 'Thechaiisready'
whole-sequence   : 'The chai is ready'
prefix-diff deltas: ['The', ' ch', 'ai', ' is', ' ready']
joined deltas    : 'The chai is ready'

This downloads the tokenizer files only, around a megabyte, not the model weights. Run against transformers 5.6.2.

'Thechaiisready'. Every space is gone. This tokenizer marks a leading space with a special symbol inside the token, and decoding one token at a time strips it as leading whitespace. Decoding tokens individually and concatenating is wrong for every SentencePiece-derived tokenizer, which is most of the Llama, Mistral and Gemma families.

The prefix-diff deltas are correct. 'The', ' ch', 'ai', ' is', ' ready' — spaces attached to the right tokens, and joining them reproduces the sentence. Decoding the full prefix each step is $O(n^2)$ over a long generation, so production detokenizers keep a small window of recent tokens and diff within it. The principle is the same: decode a span, not a single token.

What a real streaming server holds back

Three separate buffers, and each exists for its own reason:

  1. Incomplete UTF-8 bytes — until the character is whole.
  2. Possible stop-string prefixes — up to len(stop) - 1 characters, covered in stopping criteria.
  3. Tokenizer context — a few recent tokens so the decoder knows how to join.

All three are hold-backs. All three are missing from most hand-rolled streaming loops.

The wire format

Server-sent events, which is HTTP with a specific body shape:

Content-Type: text/event-stream
Cache-Control: no-cache
X-Accel-Buffering: no

data: {"choices":[{"delta":{"content":"The"}}]}

data: {"choices":[{"delta":{"content":" chai"}}]}

data: [DONE]

Each frame is data: plus one JSON object, terminated by a blank line. [DONE] is a convention rather than part of the SSE specification.

X-Accel-Buffering: no matters more than it looks. Nginx buffers proxied responses by default. A correct streaming server behind a default reverse proxy then delivers the whole answer in one lump. This is the most common "streaming does not work in production but works locally" cause.

Common mistakes

Decoding each token alone. Shown above. Spaces vanish, and non-English text breaks worse.

No client-side cancellation. If the user closes the tab, the server should stop generating. Without this you pay for tokens nobody will read and you hold a KV cache slot.

Buffering in the framework. Flask, Django and Express all need explicit streaming responses. A generator that is silently collected into a string is easy to write by accident.

Streaming a response that must be parsed as a whole. For JSON output, streaming gives the user a growing blob of syntax. Stream it only if you have a partial-JSON parser on the client, and check finish_reason before parsing.

Measuring only total time. Time to first token and inter-token latency are what streaming changes. Total time is barely affected. Log all three separately.

Try it yourself

Change CHUNKS to slice at 1 byte instead of 5. Every piece now fails the naive decode and the incremental decoder still reconstructs the sentence exactly. Then try a string of pure ASCII and watch the naive version work perfectly — which is precisely how this bug survives testing and ships.

What to learn next

Researcher — Mathematics and papers.

Why detokenization is not the inverse of tokenization

Tokenizers are not injective at the piece level. Three representative behaviours:

  • Byte-level BPE (GPT-2, Llama 3, Qwen) maps bytes into a printable alias alphabet, so a token's string may hold half a UTF-8 sequence. Only concatenation and a final byte-level decode reconstruct the text.
  • SentencePiece unigram and BPE (Llama 2, Mistral, Gemma) encode a leading space as ▁. Decoding a single token strips it as leading whitespace, so per-token decode loses word boundaries — the failure reproduced above.
  • WordPiece (BERT) marks continuation with ## and requires the opposite join rule.

The general contract is therefore: detokenize a span, not a token. Streaming detokenizers keep a rolling window of recent token ids, decode the window, and emit the difference against what has already been sent. vLLM's incremental detokenizer keeps a small prefix_offset and read_offset pair over the token list. That avoids the $O(n^2)$ cost of decoding the whole prefix each step.

Grapheme clusters, which the byte level does not solve

Correct UTF-8 decoding gives valid code points. It does not give stable user-perceived characters.

A Devanagari syllable such as त plus ै is two code points and one grapheme cluster. Emoji sequences are worse: a family emoji is several code points joined by zero-width joiners, and a flag is two regional-indicator code points.

Splitting a delta between code points of one cluster is byte-valid and visually unstable. The client renders an intermediate form, which changes when the next delta arrives. Whether to buffer to grapheme-cluster boundaries is a client decision with a latency cost, and correctness at the byte level is the server's responsibility either way.

The Unicode text-segmentation algorithm (UAX #29) defines the cluster boundaries. Python's standard library does not implement it; regex with \X does.

Latency decomposition

$$ T_{\text{total}} = T_{\text{queue}} + T_{\text{prefill}} + \sum_{t=2}^{N} T_{\text{decode}, t} $$

Streaming reveals $T_{\text{queue}} + T_{\text{prefill}}$ as time to first token and $T_{\text{decode}}$ as inter-token latency. It does not reduce either. What it changes is perceived latency. The threshold that matters is human reading speed. Roughly 4 to 6 tokens per second is readable. Beyond about 20 tokens per second, extra speed is invisible to a reader. It is not invisible to an agent consuming the output.

Report the distribution, not the mean. A p99 TTFT of eight seconds with a p50 of four hundred milliseconds is a product problem that a mean hides completely.

Transport choices

TransportDirectionReconnectNotes
SSE over HTTPserver to clientbuilt into the browser APIThe de facto standard for LLM APIs
WebSocketboth waysmanualNeeded for interruption and voice
HTTP chunked transferserver to clientnoneWhat SSE is built on
gRPC streamingboth waysbuilt inCommon between internal services

SSE won for LLM APIs because it is plain HTTP, survives proxies and CDNs, and needs no protocol upgrade. Its constraints are real: text only, one direction, and a browser limit on concurrent connections per origin over HTTP/1.1 — six, which HTTP/2 removes by multiplexing.

Interaction with the rest of the decoding stack

Speculative decoding produces several tokens in one step, so deltas arrive in bursts rather than evenly. Inter-token latency becomes bimodal, and a mean hides that. Some servers deliberately pace bursts to keep the visual rhythm even.

Stop strings require the hold-back described above. A server that streams before checking has already leaked the stop text.

Structured output streams syntactically valid prefixes that are not parseable JSON. Either buffer to completion or use a partial parser on the client.

Cancellation must propagate from a closed client connection to the scheduler, or the request keeps generating and keeps holding KV blocks. In vLLM this is an abort on the request id; without it, an abandoned long generation is pure waste.

Further reading

  • WHATWG HTML specification, Server-sent events — the transport, including reconnection semantics.
  • Unicode Standard Annex #29, Unicode Text Segmentation — grapheme cluster boundaries.
  • vLLM source, detokenize_incrementally — the reference incremental detokenizer.
  • HuggingFace transformers, TextIteratorStreamer — the simplest correct streaming implementation to read.

What to learn next

What to learn next

These follow on from what you just read.

  • Fast Attention and Long Context

    FlashAttention

    FlashAttention gives the same answer as ordinary attention while never writing the giant score table to memory, which is what makes long contexts affordable.

  • Fast Attention and Long Context

    Memory-bound vs compute-bound

    Most of what looks slow in a language model is not arithmetic, it is waiting for numbers to arrive from memory, and knowing which one you are in tells you what to fix.

  • Fast Attention and Long Context

    Multi-query attention

    Multi-query attention keeps many question-asking heads but gives them one shared set of keys and values, which shrinks the memory a model must re-read for every word it writes.