How Text Is Generated

How much memory the KV cache eats

The cache costs a fixed number of bytes per token, and multiplying that by context length and users decides how many people your GPU can serve at once.

Read these first

On this page 7
  1. Why this is the number that matters
  2. What makes the slice bigger or smaller
  3. How it works
  4. The number that surprises people
  5. Where you have already seen this
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Every token in a conversation occupies a fixed amount of graphics-card memory. That number decides how many people your server can serve at once.

Think of a hotel with a fixed number of rooms. Each guest needs one room per night. You do not need to know anything about the guest. You need the room count and the number of nights.

The model's notes work the same way. Each token books a fixed slice of memory, and holds it for the whole conversation.

The good news is that the slice size is knowable before you build anything. It comes from the model's own settings, published openly, and a few multiplications give you the answer.

Why this is the number that matters

A graphics card has a fixed amount of fast memory. On a big server card that is often eighty gigabytes.

Three things compete for it. The model's own weights go first and never move. Then working space for the calculations. Whatever is left over is for conversation notes.

So the real question is not how many users this box handles. It is how many tokens of notes fit in the leftovers.

What makes the slice bigger or smaller

Four settings from the model's configuration file decide it.

Layers. A transformer is a stack. Every layer keeps its own notes, so twice the layers means twice the memory.

Key-value heads. Each layer looks at the text through several heads at once. More heads, more notes. Modern models deliberately share heads to shrink this, which is the single biggest saving anyone has found.

Head size. How many numbers describe each token inside one head.

Number precision. Whether each number takes two bytes or one. Halving it halves the bill.

And then two things you control. Context length — how long the conversation is allowed to get. And how many conversations run at the same time.

How it works

   WHAT ONE TOKEN COSTS

     start with two, one for keys and one for values
     multiply by the number of layers
     multiply by the number of key-value heads
     multiply by the head size
     multiply by the bytes per number

   WHAT THE WHOLE SERVER COSTS

     that per-token cost
     once for every token in the conversation
     once for every conversation running at the same time

That is the whole recipe, and it is exact. There is no hidden overhead worth worrying about.

The number that surprises people

For a popular eight-billion-parameter model, one token of notes costs about one hundred and twenty-eight kilobytes.

That sounds tiny. A conversation of a hundred and twenty-eight thousand tokens costs sixteen gigabytes. For one user.

The weights of that same model take about fifteen gigabytes. So a single long conversation can cost more memory than the entire model.

This is the fact that catches every team the first time they run out of memory in production.

Where you have already seen this

  • A chat tool that gets noticeably more expensive per message in very long threads.
  • A provider charging more for a long-context version of the same model.
  • A self-hosted model that runs fine alone and crashes with four users.
  • A service that quietly truncates the oldest part of your conversation.

Remember this

  • Cache memory is bytes per token times context length times users. Nothing more.
  • Bytes per token comes from the model's config file, and you can read it today.
  • On long contexts the notes can outgrow the model itself.

What to learn next

Developer — Code and libraries.

Setup

bash
# no libraries needed
python3 --version

This is arithmetic, not machine learning. Every number below comes from a config.json published on the HuggingFace Hub.

The calculator

kv_memory.py
def kv_bytes(layers, kv_heads, head_dim, tokens, batch=1, bytes_per_number=2):
    """Bytes held by the KV cache. The 2 is for the K and the V."""
    return 2 * layers * kv_heads * head_dim * tokens * batch * bytes_per_number


def gib(n):
    return n / 1024 ** 3


# Numbers read straight out of each model's config.json on the Hub.
MODELS = {
    #                    layers  kv_heads  head_dim
    "Qwen3-0.6B":          (28,     8,       128),
    "Llama-3.1-8B":        (32,     8,       128),
    "Qwen3-32B":           (64,     8,       128),
}

print(f"{'model':<16}{'per token':>12}{'4k ctx':>10}{'32k ctx':>10}{'128k ctx':>11}")
for name, (L, H, Dh) in MODELS.items():
    per_tok = kv_bytes(L, H, Dh, 1)
    row = f"{name:<16}{per_tok / 1024:>10.0f} KB"
    for ctx in (4096, 32768, 131072):
        row += f"{gib(kv_bytes(L, H, Dh, ctx)):>9.2f}G"
    print(row)

print("\nGrouped-query attention is the whole reason these numbers are survivable.")
L, Dh, ctx = 32, 128, 32768
for label, heads in (("multi-head (32 KV heads)", 32), ("grouped-query (8 KV heads)", 8)):
    print(f"  {label:<28}{gib(kv_bytes(L, heads, Dh, ctx)):>7.2f} GiB at 32k tokens")

print("\nWhat one 80 GiB GPU can actually hold, Llama-3.1-8B in bf16:")
weights = 8.03e9 * 2                       # 8.03 B parameters at 2 bytes each
budget = 80 * 1024 ** 3 - weights - 4 * 1024 ** 3   # minus weights, minus working room
per_seq = kv_bytes(32, 8, 128, 4096)
print(f"  weights                : {gib(weights):.1f} GiB")
print(f"  left for the KV cache  : {gib(budget):.1f} GiB")
print(f"  one 4k-token sequence  : {gib(per_seq) * 1024:.0f} MiB")
print(f"  concurrent 4k sequences: {int(budget // per_seq)}")
Output
model              per token    4k ctx   32k ctx   128k ctx
Qwen3-0.6B             112 KB     0.44G     3.50G    14.00G
Llama-3.1-8B           128 KB     0.50G     4.00G    16.00G
Qwen3-32B              256 KB     1.00G     8.00G    32.00G

Grouped-query attention is the whole reason these numbers are survivable.
  multi-head (32 KV heads)      16.00 GiB at 32k tokens
  grouped-query (8 KV heads)     4.00 GiB at 32k tokens

What one 80 GiB GPU can actually hold, Llama-3.1-8B in bf16:
  weights                : 15.0 GiB
  left for the KV cache  : 61.0 GiB
  one 4k-token sequence  : 512 MiB
  concurrent 4k sequences: 122

Reading the output

Qwen3-0.6B costs 112 KB per token, and Llama-3.1-8B costs 128 KB. The models differ in size by more than a factor of ten. Their cache costs differ by fourteen percent. Cache cost tracks layer count and KV heads, not parameter count, and this catches people out constantly. A small model with deep layers and many KV heads can have a heavier cache than a larger model with aggressive grouping.

Qwen3-32B at 128k context needs 32 GiB of cache for one sequence. Its weights in bf16 are roughly 65 GiB. One user at full context adds half the model again.

16.00 GiB against 4.00 GiB. Grouped-query attention divides KV heads by four here and divides the cache by four. This is not a rounding-error optimisation. It is the difference between a serviceable model and an unservable one, which is why every current model uses it.

122 concurrent 4k sequences. That is the honest capacity number for a single H100-class card, before you account for fragmentation, which PagedAttention exists to fix. Serving frameworks report this as the number of KV blocks and it is the first log line worth reading.

Getting the numbers for any model

python
from transformers import AutoConfig

cfg = AutoConfig.from_pretrained("Qwen/Qwen3-0.6B")
layers = cfg.num_hidden_layers
kv_heads = getattr(cfg, "num_key_value_heads", cfg.num_attention_heads)
head_dim = getattr(cfg, "head_dim", None) or cfg.hidden_size // cfg.num_attention_heads
print(layers, kv_heads, head_dim)

Note the two getattr calls, and do not skip them. num_key_value_heads is absent on older multi-head models, where it equals num_attention_heads. head_dim is absent on many Llama-family configs, where it is hidden_size // num_attention_heads. Assuming either field exists is the most common bug in home-made capacity calculators.

This downloads a small JSON file, not the model weights.

What actually consumes your GPU

Four things, in decreasing order of predictability:

  1. Weights. Parameters times bytes per parameter. Fixed and exact.
  2. KV cache. The formula above. Grows with traffic.
  3. Activations. Transient working memory during a forward pass. Grows with batch size and prefill chunk size.
  4. Fragmentation and allocator overhead. Real, awkward, and the reason production frameworks manage memory themselves.

Serving frameworks ask you for a fraction of the card to use, and carve the KV pool out of whatever is left after weights. Set that fraction too high and you get an out-of-memory error during a traffic spike, not at start-up. That is the worst possible time to find out.

Levers that actually work

Grouped-query attention. Not a runtime choice — it is baked into the architecture. It is a model-selection criterion. Check num_key_value_heads before you commit to a model.

KV cache quantisation. Storing the cache in fp8 rather than bf16 halves it. vLLM exposes this as kv_cache_dtype="fp8". Quality impact is small but real, and it should be measured on your own task, not taken from a blog post.

Lower the advertised maximum context. Serving frameworks preallocate against max_model_len. Dropping it from 128k to 8k when your product never needs more is the largest single saving available, and it costs nothing.

Prefix sharing. Identical system prompts across users can point at the same physical blocks. Free, if your prompt layout keeps the shared part first.

More cards. Tensor parallelism splits KV heads across GPUs, so cache per card falls with the parallel degree. It also adds communication cost per token.

Common mistakes

Using parameter count as a proxy for cache size. Shown false above. Read the config.

Forgetting the factor of two for K and V. Half of every home-made estimate on the internet is wrong by exactly this.

Sizing against average context length. The cache is held for the whole request, and your capacity is set by the peak, not the mean. Size against the maximum you allow, then reduce the maximum.

Confusing GiB and GB. A card advertised as 80 GB is 80 GiB of usable HBM in NVIDIA's own reporting. Mixing the two conventions produces a seven-percent error, in exactly the direction that breaks you.

Try it yourself

Add an old multi-head model to the table, using kv_heads = num_attention_heads. Compute how many concurrent 4k sequences it supports on the same 80 GiB card. Then explain the result to somebody as a per-user cost in rupees or dollars per hour of GPU rental.

What to learn next

Researcher — Mathematics and papers.

The exact expression

For a dense transformer with $L$ layers, $H_{kv}$ key-value heads, head dimension $d_h$, batch $B$, sequence length $S$, and $b$ bytes per stored element:

$$ M_{\text{KV}} = 2 \cdot L \cdot H_{kv} \cdot d_h \cdot S \cdot B \cdot b $$

The leading $2$ counts keys and values. Under multi-head attention $H_{kv} = H$ and $H \cdot d_h = d_{\text{model}}$, which reduces the expression to $2 L d_{\text{model}} S B b$. That is the form most often quoted, and it stops being correct the moment grouped-query attention is used.

Define the grouping ratio $g = H / H_{kv}$. Then

$$ M_{\text{KV}} = \frac{2 L d_{\text{model}} S B b}{g} $$

Llama 3 and Qwen 3 use $g = 4$ to $g = 8$ at the larger sizes.

Cache traffic, which is what actually limits decode

Memory footprint sets capacity. Memory bandwidth sets speed, and these are different constraints.

Each decode step reads the entire cache for the sequences in the batch. Per step:

$$ \text{bytes read} \approx M_{\text{KV}} \quad\text{(the whole cache)} $$

so time per output token has a floor of $M_{\text{KV}} / \text{BW}$. On an H100 at $3.35$ TB/s, a batch holding 40 GiB of cache cannot decode faster than about $12.8$ ms per step. Arithmetic speed is irrelevant to that floor. Attention kernels are memory-bound at decode time for exactly this reason. FlashAttention-style tiling helps by avoiding materialising the score matrix, not by reading fewer cache bytes.

The consequence is a genuine ceiling. Growing batch size raises throughput until cache traffic saturates bandwidth, and beyond that point every extra user slows everyone down proportionally.

Architectural responses

ApproachCache per tokenCost
Multi-head (MHA)$2 L H d_h b$baseline
Multi-query (MQA)$2 L d_h b$quality loss at scale
Grouped-query (GQA)$2 L (H/g) d_h b$small quality loss, universal in practice
Multi-head latent (MLA)$2 L d_c b$, $d_c \ll H d_h$extra arithmetic to reconstruct K and V
Sliding window$2 L H_{kv} d_h w b$, window $w$lossy beyond the window
Cross-layer sharingdivided by layer group sizequality loss, active research

MLA (DeepSeek-V2) compresses K and V into a shared latent of dimension $d_c$ and reconstructs per head during attention. Reported cache reduction is on the order of $90\%$ against MHA at comparable quality. The reconstruction can be folded into the query and output projections, so the extra arithmetic is largely free during decode, which is memory-bound anyway.

Cache quantisation has an asymmetric error profile. Keys carry outlier channels that quantise poorly; values are better behaved. Per-channel scaling for keys and per-token scaling for values is the usual compromise (Liu et al., KIVI, 2024, arxiv.org/abs/2402.02750).

Eviction and compression. H2O (Zhang et al., 2023, arxiv.org/abs/2306.14048) keeps a "heavy hitter" subset of tokens, chosen by accumulated attention mass. SnapKV and similar methods compress the prompt's cache after prefill. All of these are lossy, and their reported quality preservation is heavily benchmark-dependent. Treat published retention numbers as an upper bound on what you will see.

Capacity planning, done honestly

$$ B_{\max} = \left\lfloor \frac{M_{\text{GPU}} - M_{\text{weights}} - M_{\text{activations}}}{M_{\text{KV per sequence}}} \right\rfloor $$

Three warnings about applying this.

$M_{\text{KV per sequence}}$ should use the maximum allowed context, not the observed mean, unless your scheduler can preempt. With paged allocation and preemption, planning against a percentile of the length distribution becomes defensible. Preemption itself has a cost, since a preempted sequence is either recomputed or swapped to host memory.

$M_{\text{activations}}$ scales with the prefill chunk size, so it is a scheduler parameter rather than a constant.

And the whole calculation assumes a steady state that real traffic does not have. Size for the peak, then measure preemption rate as your capacity signal.

Papers

What to learn next

What to learn next

These follow on from what you just read.

  • How Text Is Generated

    PagedAttention

    Serving systems stopped reserving one long contiguous slab of memory per conversation and started handing out small fixed-size blocks on demand, which recovered most of the memory that used to be wasted.

  • How Text Is Generated

    Top-k and nucleus sampling

    Two ways to throw away the model's worst options before picking a word — keep a fixed number of them, or keep enough to cover a fixed share of the model's confidence.

  • How Text Is Generated

    Min-p and typical sampling

    Two truncation methods that decide the shortlist from how confident the model is at that exact step, instead of from a fixed count or a fixed share of probability.