How Text Is Generated

Prefill and decode

A model reads your entire prompt in one fast pass, then writes the answer one token at a time in a slow loop. Almost every delay in text generation comes from that split.

On this page 7
  1. Why the split exists at all
  2. How it looks
  3. What this means for you as a user
  4. Where you have already seen this
  5. What is honestly hard here
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Reading your prompt and writing the answer are two different jobs, and the model does them in two different ways.

Think about replying to a long letter by hand. You read the whole page in one sitting, taking it in quickly. Then you write your reply word by word. Before each new word, your eyes flick back over what you have already written.

Reading is fast because you do it all at once. Writing is slow because every word waits for the one before it.

A language model works the same way. The reading phase is called prefill — pushing your whole prompt through the model in a single pass. The writing phase is called decode — producing one token at a time. A token is a small chunk of text, usually a word or part of a word.

Why the split exists at all

A model has one skill. Given some text, it guesses what comes next.

That is the whole trick. There is no separate answering machine hidden inside it.

So to write ten words, the model has to run ten times. Each run needs the previous word, because that word is part of the text it reads next time.

Your prompt is different. Every word of it already exists before the model starts. So the model can look at all of them together, in one go.

How it looks

   PREFILL  (one pass, all your words at once)

     "what   is   the   capital   of"
        |     |     |      |       |
        +-----+-----+------+-------+
                    |
                    v
              first new token: India


   DECODE  (a loop, one token per pass)

     ... of India                  ->  has
     ... of India has              ->  many
     ... of India has many         ->  languages

Notice the shape. Prefill is one wide step. Decode is a long thin ladder.

What this means for you as a user

The pause before the first word appears is prefill. Engineers call it time to first token — the delay between pressing send and seeing anything at all.

The speed at which words then stream out is decode. That is a steady drip, and it stays roughly the same speed whether your answer is short or long.

A very long prompt makes the first pause longer. A very long answer makes the streaming part longer. These are two separate complaints with two separate fixes.

Where you have already seen this

  • ChatGPT sitting still for a moment, then streaming text at a readable pace.
  • A coding assistant taking longer to start when your file is large.
  • A voice assistant that begins speaking before it has finished thinking.
  • A summarise-this-PDF tool that pauses hard on a hundred-page document.

What is honestly hard here

Decode feels like it should be fast. Each step does far less work than prefill does in total.

It is slow for a reason that surprises everyone the first time. For every single token it writes, the model fetches all of its own memory. That means every number it stored during training. Fetching is the bottleneck, not thinking.

This part confuses almost everyone at first. Read it twice. Most of this section exists because of it.

Remember this

  • Prefill reads your whole prompt in one pass, so it is wide and quick.
  • Decode writes one token per pass, so it is a slow loop you cannot skip.
  • Long prompts hurt the first pause; long answers hurt the streaming speed.

What to learn next

  • The KV cache — the fix for the redoing-everything problem you saw above.
  • How LLMs work — the model these two phases are running.
  • vLLM — a server built around exactly this split.

Developer — Code and libraries.

Setup

bash
pip install numpy

The whole point fits in one file with no model download. We build a toy attention layer, run a prompt through it once, then generate from it token by token, and count the work.

Watching the two phases happen

prefill_decode.py
import numpy as np

rng = np.random.default_rng(0)
V, D = 12, 8                                  # 12-word vocabulary, 8 numbers per token

E  = rng.normal(size=(V, D)) * 0.5            # embedding table
Wq = rng.normal(size=(D, D)) * 0.5
Wk = rng.normal(size=(D, D)) * 0.5
Wv = rng.normal(size=(D, D)) * 0.5
Wo = rng.normal(size=(D, V)) * 0.5

token_rows = 0        # how many token-rows we push through the weight matrices


def softmax(z):
    z = z - z.max(axis=-1, keepdims=True)
    e = np.exp(z)
    return e / e.sum(axis=-1, keepdims=True)


def forward(tokens):
    """Run a whole sequence through one attention layer. Logits for EVERY position."""
    global token_rows
    x = E[list(tokens)]                       # (n, D)
    token_rows += x.shape[0]                  # this is the work that costs money
    q, k, v = x @ Wq, x @ Wk, x @ Wv
    n = len(tokens)
    mask = np.triu(np.full((n, n), -np.inf), 1)   # no position may look at the future
    w = softmax(q @ k.T / np.sqrt(D) + mask)
    return (w @ v) @ Wo                       # (n, V)


prompt = [3, 7, 1, 9, 4, 2]                   # a six-token "prompt"

# ---- PREFILL: the whole prompt in one pass -------------------------------
logits = forward(prompt)
print("PREFILL")
print("  prompt length        :", len(prompt))
print("  logits shape         :", logits.shape, "  (one row per prompt position)")
print("  rows we actually use : 1                (only the last row predicts token 7)")
print("  token-rows computed  :", token_rows)

# ---- DECODE: one token at a time ----------------------------------------
print("\nDECODE (no cache yet, so each step re-runs the whole prefix)")
seq = list(prompt)
before = token_rows
for step in range(4):
    logits = forward(seq)
    nxt = int(logits[-1].argmax())            # greedy: take the highest-scoring token
    seq.append(nxt)
    print(f"  step {step + 1}: context {len(seq) - 1:>2} tokens"
          f" -> new token {nxt:>2}   (token-rows this step: {len(seq) - 1})")

print("\nSUMMARY")
print("  tokens generated     : 4")
print("  token-rows in prefill:", before)
print("  token-rows in decode :", token_rows - before)
print("  final sequence       :", seq)
Output
PREFILL
  prompt length        : 6
  logits shape         : (6, 12)   (one row per prompt position)
  rows we actually use : 1                (only the last row predicts token 7)
  token-rows computed  : 6

DECODE (no cache yet, so each step re-runs the whole prefix)
  step 1: context  6 tokens -> new token  9   (token-rows this step: 6)
  step 2: context  7 tokens -> new token  1   (token-rows this step: 7)
  step 3: context  8 tokens -> new token  3   (token-rows this step: 8)
  step 4: context  9 tokens -> new token  1   (token-rows this step: 9)

SUMMARY
  tokens generated     : 4
  token-rows in prefill: 6
  token-rows in decode : 30
  final sequence       : [3, 7, 1, 9, 4, 2, 9, 1, 3, 1]

Reading the output line by line

Prefill produced 6 rows of logits and used 1. A logit is the raw score a model gives each vocabulary word before those scores become probabilities. Every prompt position gets its own row of scores, but only the last row predicts the token that comes next. The other five rows are not wasted during training — they are the training signal. At inference time they are discarded.

Decode did four passes and grew each time. Six token-rows, then seven, then eight, then nine. The context is one token longer every step, so the same work gets redone plus a little more.

Thirty token-rows to produce four tokens. Prefill did six rows for six tokens. Decode did thirty rows for four. That ratio is the problem the rest of this section attacks, starting with the KV cache.

The output degenerates into a repeated token 1. An untrained random model has no reason to say anything sensible. Notice that it also gets stuck, which previews repetition penalties.

Why decode is slow, stated properly

Prefill and decode run the same weights. The difference is how much work each fetched byte of weights supports.

In prefill, one read of a weight matrix serves every token in your prompt at once. Two thousand prompt tokens ride on a single fetch.

In decode at batch size one, one read of the same matrix serves exactly one token. The arithmetic units sit idle waiting for memory. This is called being memory-bandwidth bound — limited by how fast numbers move, not by how fast they multiply.

That is why decode speed barely improves on a faster GPU with the same memory bandwidth. It is also why serving many users at once is nearly free per extra user.

Common mistakes

Benchmarking with a one-token prompt and calling it inference speed. You measured decode with an empty cache. Real workloads have prompts of hundreds or thousands of tokens, and prefill dominates time to first token. Report the two numbers separately, always.

Assuming a bigger batch makes each user's stream faster. It does not. Batching raises total throughput and leaves each user's tokens-per-second roughly unchanged, or slightly worse. Throughput and latency are different metrics and they trade against each other.

Forgetting that the prompt is charged too. Providers bill input tokens and output tokens separately, at different rates, because prefill and decode cost different amounts. A long system prompt repeated on every request is a real monthly bill.

Timing the first token on a cold server. The first request after start-up pays for weight loading and kernel warm-up. Discard the first few runs before measuring anything.

Try it yourself

Change prompt to twenty tokens and generate two. Then change it to two tokens and generate twenty. Total token-rows is what a GPU actually pays for, and the two cases differ wildly. Predict which is worse before you run it.

What to learn next

  • The KV cache — the fix for the redoing-everything problem you saw above.
  • How LLMs work — the model these two phases are running.
  • vLLM — a server built around exactly this split.

Researcher — Mathematics and papers.

The two phases as arithmetic

Let $P$ be the number of model parameters, $N$ the prompt length, $B$ the batch size, and $b$ the bytes per parameter (2 for bf16).

A forward pass through a dense transformer costs roughly $2$ FLOPs per parameter per token, since each weight takes part in one multiply and one add.

$$ C_{\text{prefill}} \approx 2 P N B \ \text{FLOPs}, \qquad C_{\text{decode-step}} \approx 2 P B \ \text{FLOPs} $$

Both phases must read all $P b$ bytes of weights from high-bandwidth memory. Arithmetic intensity is FLOPs performed per byte moved:

$$ I_{\text{prefill}} \approx \frac{2PNB}{Pb} = \frac{2NB}{b}, \qquad I_{\text{decode}} \approx \frac{2B}{b} $$

With $b = 2$, prefill intensity is $\approx NB$ and decode intensity is $\approx B$. An H100 SXM delivers about $989$ dense BF16 TFLOP/s against $3.35$ TB/s of HBM3 bandwidth, giving a ridge point near

$$ \frac{989 \times 10^{12}}{3.35 \times 10^{12}} \approx 295 \ \text{FLOP/byte} $$

So decode at batch 1 runs at roughly $1/295$ of peak arithmetic throughput. Prefill crosses the ridge with a few hundred tokens. This single ratio explains most of LLM serving economics.

Attention scales differently from the weights

The parameter term above ignores attention scores. Per layer, with hidden size $d$ and sequence length $n$:

  • Prefill attention: $O(n^2 d)$ FLOPs — quadratic, and it eventually dominates the linear weight term.
  • Decode attention for one token: $O(n d)$ FLOPs, but it must read the entire KV cache, which is $O(n)$ bytes. Intensity is $O(1)$ regardless of batch size, because each sequence owns its cache.

That last point matters. Batching amortises weight reads across sequences but does not amortise KV cache reads. As context grows, attention becomes the memory bottleneck even at large batch, which is what motivates PagedAttention and the whole efficient-attention literature.

The metrics that are actually reported

  • TTFT — time to first token, dominated by prefill and by queueing.
  • TPOT or ITL — time per output token, or inter-token latency, dominated by decode.
  • End-to-end latency $\approx \text{TTFT} + \text{TPOT} \times (\text{output tokens} - 1)$.
  • Throughput — total tokens per second across all concurrent requests.

Reporting a single tokens-per-second figure without saying which of these it is makes a benchmark uninterpretable. Most published comparisons fail this test.

Scheduling, where the real engineering lives

Static batching makes every request wait for the slowest one in its batch. Continuous batching (Yu et al., Orca, OSDI 2022) schedules at the granularity of one iteration instead: finished sequences leave the batch and new ones join immediately. This alone is worth several times the throughput on realistic traffic.

It creates a new problem. A long prefill occupies the GPU for one very long iteration, stalling every decode in flight, which users mid-conversation experience as a latency spike.

Chunked prefill (Agrawal et al., Sarathi-Serve, OSDI 2024) splits a long prompt into fixed-size chunks and packs each chunk into a hybrid batch alongside ongoing decodes. Decode steps are memory-bound and leave arithmetic units idle; prefill chunks fill exactly that gap. The result is stall-free scheduling with better utilisation than either phase reaches alone.

Both vLLM and SGLang enable chunked prefill by default for long prompts. Chunk size is the knob: larger chunks favour throughput, smaller chunks favour inter-token latency for streams already in flight.

Prefix caching

Prefill is redundant across requests that share a prefix — a system prompt, a few-shot block, a document being asked about repeatedly. Caching the KV state of shared prefixes turns repeated prefill into a lookup.

SGLang's RadixAttention (Zheng et al., 2024) organises the cache as a radix tree over token sequences, with LRU eviction. Prefix sharing then happens automatically, rather than through manual cache keys. vLLM exposes the same idea as automatic prefix caching.

The practical consequence for prompt design: keep the stable part of your prompt at the front and the variable part at the back. Reversing that order defeats prefix caching completely, and it is a common and expensive mistake.

Papers

  • Yu et al., Orca: A Distributed Serving System for Transformer-Based Generative Models, OSDI 2022 — continuous batching.
  • Kwon et al., Efficient Memory Management for LLM Serving with PagedAttention, SOSP 2023 — arxiv.org/abs/2309.06180
  • Agrawal et al., Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve, OSDI 2024 — arxiv.org/abs/2403.02310
  • Zheng et al., SGLang: Efficient Execution of Structured Language Model Programs, 2024 — arxiv.org/abs/2312.07104
  • Dao et al., FlashAttention, 2022 — arxiv.org/abs/2205.14135

What to learn next

  • The KV cache — the fix for the redoing-everything problem you saw above.
  • How LLMs work — the model these two phases are running.
  • vLLM — a server built around exactly this split.

What to learn next

These follow on from what you just read.

  • How Text Is Generated

    The KV cache

    Instead of re-reading the whole conversation for every new word, the model keeps a running summary of what each earlier token contributed and appends one row per new token.

  • How Text Is Generated

    How much memory the KV cache eats

    The cache costs a fixed number of bytes per token, and multiplying that by context length and users decides how many people your GPU can serve at once.

  • How Text Is Generated

    PagedAttention

    Serving systems stopped reserving one long contiguous slab of memory per conversation and started handing out small fixed-size blocks on demand, which recovered most of the memory that used to be wasted.