KV cache
In one sentence The KV cache stores each token's attention keys and values so the model never recomputes them, making generation fast but memory-hungry.
Updated
The KV cache saves the attention keys and values computed for every previous token, so generating each new token reuses them instead of recomputing everything.
Imagine an accountant totalling a long ledger. A new entry arrives. The foolish method: re-add all 5,000 rows from the top. The sane method: keep the running work on the desk and process only the new row. The KV cache is that kept work for a transformer.
The specifics: in attention, every token produces a key ("what I contain") and a value ("what I contribute"). When generating token 5,001, the model needs the keys and values of all 5,000 previous tokens — and those never change. So they are computed once and cached. Each step then computes only the new token's query, key and value, attends against the cache, and appends to it.
The consequences run through the whole economics of LLM serving:
speed : generation cost per token stays roughly constant instead of growing
memory : cache grows with every token, per layer, per attention head
— long chats can consume gigabytes per user
first token vs rest: the initial prompt must be processed in full ("prefill"),
which is why the first token takes longestThat memory pressure explains a family of modern designs: grouped-query-attention shrinks the cache by sharing keys and values across heads, cache quantization stores it in fewer bits, and vLLM's PagedAttention manages it like virtual memory.
Where to go next
- Full lesson: vLLM
- Related terms: attention, context-window, inference, quantization