How it works

How GPU inference works

Serving an LLM means one parallel pass over your prompt, then a one-token-at-a-time loop — kept fast by a cache, kept affordable by batching many users onto one GPU.

On this page 7
  1. The pipeline at a glance
  2. Stage 1 — tokenize on the CPU
  3. Stage 2 — prefill: read the whole prompt at once
  4. Stage 3 — decode: one token per pass
  5. Stage 4 — batching: sharing the pantry trip
  6. Stage 5 — stream the tokens out
  7. Which lessons teach each stage

Training gets the glory, but every ChatGPT reply you have ever read came from the other half of the field: inference — running a finished model to produce output. Serving one model to millions of users is its own engineering discipline. This page follows one request through a GPU, and shows where the speed and the cost actually come from.

The pipeline at a glance

 request: "Summarise this article: ..."
        |
        v
 [1. tokenize]      text -> token IDs
        |
        v
 [2. prefill]       whole prompt processed in ONE parallel pass
        |           (GPU at full power; writes the KV cache)
        v
 [3. decode loop]   ONE new token per pass
        |    ^      (memory-bound; reads the KV cache)
        |    |__ append token, repeat
        v
 [4. batching]      many users' loops share the same GPU
        |
        v
 [5. stream]        tokens sent back as they are made

Stage 1 — tokenize on the CPU

The text is chopped into tokens — small chunks of text — and mapped to ID numbers. Cheap, quick, done on the CPU. The IDs are shipped to the GPU, where everything expensive happens.

Stage 2 — prefill: read the whole prompt at once

The model processes your entire prompt in a single parallel pass, called prefill. Every token is handled at the same time, which is exactly what a GPU is for: thousands of small cores doing the same arithmetic on different data at once.

Prefill produces two things. The first generated token — and, more importantly, the KV cache: stored attention values for every prompt token. Attention requires each new token to look back at all previous ones. The cache keeps those look-back values so they are computed once, not recomputed for every future token.

The pause before ChatGPT starts answering a long prompt? That is prefill running. Servers measure it as time-to-first-token.

Stage 3 — decode: one token per pass

Now generation begins, and the rhythm changes completely. Each pass produces exactly one token, which is appended, and the model runs again. A 500-token answer means 500 sequential passes. No parallel trick removes this: token 2 cannot be computed until token 1 exists.

Here is the counter-intuitive part. Each decode pass does little arithmetic, but must stream all the model's weights — tens of gigabytes — from GPU memory through the compute units, for one token. Decode speed is limited by memory bandwidth, not by calculation. The GPU is a chef with lightning hands whose every dish requires fetching the entire pantry.

That explains quantization: storing weights in fewer bits — 8, or 4, instead of 16. Half the bytes to move per pass means roughly double the tokens per second, and a model that fits on a smaller, cheaper GPU. A small accuracy cost buys a large speed and cost win.

Stage 4 — batching: sharing the pantry trip

If moving the weights is the bottleneck, serve many users per trip. Batching runs one pass that decodes the next token for dozens of requests at once — the weights move through the chip once, and every request in the batch benefits.

Naive batching waits for the whole batch to finish, so one user's 2,000-token essay holds everyone hostage. Modern servers like vLLM use continuous batching: finished requests exit the batch immediately and waiting ones slot in, every single pass. This one scheduling idea multiplied real-world GPU throughput several times over, and is a large part of why API prices keep falling.

Stage 5 — stream the tokens out

Tokens are sent to the user as they are produced, which is why the answer types itself out. Streaming does not make generation faster. It makes waiting honest: you read at the speed the model writes, instead of staring at a spinner until the end.

The KV cache has one more trick here: if many requests share a prefix — the same system prompt, say — its cached values can be computed once and reused across requests. That is the mechanism behind "prompt caching" discounts on API bills.

Which lessons teach each stage