How much memory the KV cache eats
The cache costs a fixed number of bytes per token, and multiplying that by context length and users decides how many people your GPU can serve at once.
- 13 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Every token in a conversation occupies a fixed amount of graphics-card memory. That number decides how many people your server can serve at once.
Think of a hotel with a fixed number of rooms. Each guest needs one room per night. You do not need to know anything about the guest. You need the room count and the number of nights.
The model's notes work the same way. Each token books a fixed slice of memory, and holds it for the whole conversation.
The good news is that the slice size is knowable before you build anything. It comes from the model's own settings, published openly, and a few multiplications give you the answer.
Why this is the number that matters
A graphics card has a fixed amount of fast memory. On a big server card that is often eighty gigabytes.
Three things compete for it. The model's own weights go first and never move. Then working space for the calculations. Whatever is left over is for conversation notes.
So the real question is not how many users this box handles. It is how many tokens of notes fit in the leftovers.
What makes the slice bigger or smaller
Four settings from the model's configuration file decide it.
Layers. A transformer is a stack. Every layer keeps its own notes, so twice the layers means twice the memory.
Key-value heads. Each layer looks at the text through several heads at once. More heads, more notes. Modern models deliberately share heads to shrink this, which is the single biggest saving anyone has found.
Head size. How many numbers describe each token inside one head.
Number precision. Whether each number takes two bytes or one. Halving it halves the bill.
And then two things you control. Context length — how long the conversation is allowed to get. And how many conversations run at the same time.
How it works
WHAT ONE TOKEN COSTS
start with two, one for keys and one for values
multiply by the number of layers
multiply by the number of key-value heads
multiply by the head size
multiply by the bytes per number
WHAT THE WHOLE SERVER COSTS
that per-token cost
once for every token in the conversation
once for every conversation running at the same timeThat is the whole recipe, and it is exact. There is no hidden overhead worth worrying about.
The number that surprises people
For a popular eight-billion-parameter model, one token of notes costs about one hundred and twenty-eight kilobytes.
That sounds tiny. A conversation of a hundred and twenty-eight thousand tokens costs sixteen gigabytes. For one user.
The weights of that same model take about fifteen gigabytes. So a single long conversation can cost more memory than the entire model.
This is the fact that catches every team the first time they run out of memory in production.
Where you have already seen this
- A chat tool that gets noticeably more expensive per message in very long threads.
- A provider charging more for a long-context version of the same model.
- A self-hosted model that runs fine alone and crashes with four users.
- A service that quietly truncates the oldest part of your conversation.
Remember this
- Cache memory is bytes per token times context length times users. Nothing more.
- Bytes per token comes from the model's config file, and you can read it today.
- On long contexts the notes can outgrow the model itself.
What to learn next
- PagedAttention — how servers stop wasting most of the memory you calculated above.
- Quantization in practice — the same bytes-per-number lever, applied to weights.
- Latency and throughput — the two numbers this capacity calculation trades between.
Developer — Code and libraries.
Setup
# no libraries needed
python3 --versionThis is arithmetic, not machine learning. Every number below comes from a config.json published on the HuggingFace Hub.
The calculator
def kv_bytes(layers, kv_heads, head_dim, tokens, batch=1, bytes_per_number=2):
"""Bytes held by the KV cache. The 2 is for the K and the V."""
return 2 * layers * kv_heads * head_dim * tokens * batch * bytes_per_number
def gib(n):
return n / 1024 ** 3
# Numbers read straight out of each model's config.json on the Hub.
MODELS = {
# layers kv_heads head_dim
"Qwen3-0.6B": (28, 8, 128),
"Llama-3.1-8B": (32, 8, 128),
"Qwen3-32B": (64, 8, 128),
}
print(f"{'model':<16}{'per token':>12}{'4k ctx':>10}{'32k ctx':>10}{'128k ctx':>11}")
for name, (L, H, Dh) in MODELS.items():
per_tok = kv_bytes(L, H, Dh, 1)
row = f"{name:<16}{per_tok / 1024:>10.0f} KB"
for ctx in (4096, 32768, 131072):
row += f"{gib(kv_bytes(L, H, Dh, ctx)):>9.2f}G"
print(row)
print("\nGrouped-query attention is the whole reason these numbers are survivable.")
L, Dh, ctx = 32, 128, 32768
for label, heads in (("multi-head (32 KV heads)", 32), ("grouped-query (8 KV heads)", 8)):
print(f" {label:<28}{gib(kv_bytes(L, heads, Dh, ctx)):>7.2f} GiB at 32k tokens")
print("\nWhat one 80 GiB GPU can actually hold, Llama-3.1-8B in bf16:")
weights = 8.03e9 * 2 # 8.03 B parameters at 2 bytes each
budget = 80 * 1024 ** 3 - weights - 4 * 1024 ** 3 # minus weights, minus working room
per_seq = kv_bytes(32, 8, 128, 4096)
print(f" weights : {gib(weights):.1f} GiB")
print(f" left for the KV cache : {gib(budget):.1f} GiB")
print(f" one 4k-token sequence : {gib(per_seq) * 1024:.0f} MiB")
print(f" concurrent 4k sequences: {int(budget // per_seq)}")model per token 4k ctx 32k ctx 128k ctx Qwen3-0.6B 112 KB 0.44G 3.50G 14.00G Llama-3.1-8B 128 KB 0.50G 4.00G 16.00G Qwen3-32B 256 KB 1.00G 8.00G 32.00G Grouped-query attention is the whole reason these numbers are survivable. multi-head (32 KV heads) 16.00 GiB at 32k tokens grouped-query (8 KV heads) 4.00 GiB at 32k tokens What one 80 GiB GPU can actually hold, Llama-3.1-8B in bf16: weights : 15.0 GiB left for the KV cache : 61.0 GiB one 4k-token sequence : 512 MiB concurrent 4k sequences: 122
Reading the output
Qwen3-0.6B costs 112 KB per token, and Llama-3.1-8B costs 128 KB. The models differ in size by more than a factor of ten. Their cache costs differ by fourteen percent. Cache cost tracks layer count and KV heads, not parameter count, and this catches people out constantly. A small model with deep layers and many KV heads can have a heavier cache than a larger model with aggressive grouping.
Qwen3-32B at 128k context needs 32 GiB of cache for one sequence. Its weights in bf16 are roughly 65 GiB. One user at full context adds half the model again.
16.00 GiB against 4.00 GiB. Grouped-query attention divides KV heads by four here and divides the cache by four. This is not a rounding-error optimisation. It is the difference between a serviceable model and an unservable one, which is why every current model uses it.
122 concurrent 4k sequences. That is the honest capacity number for a single H100-class card, before you account for fragmentation, which PagedAttention exists to fix. Serving frameworks report this as the number of KV blocks and it is the first log line worth reading.
Getting the numbers for any model
from transformers import AutoConfig
cfg = AutoConfig.from_pretrained("Qwen/Qwen3-0.6B")
layers = cfg.num_hidden_layers
kv_heads = getattr(cfg, "num_key_value_heads", cfg.num_attention_heads)
head_dim = getattr(cfg, "head_dim", None) or cfg.hidden_size // cfg.num_attention_heads
print(layers, kv_heads, head_dim)Note the two getattr calls, and do not skip them. num_key_value_heads is absent on older multi-head models, where it equals num_attention_heads. head_dim is absent on many Llama-family configs, where it is hidden_size // num_attention_heads. Assuming either field exists is the most common bug in home-made capacity calculators.
This downloads a small JSON file, not the model weights.
What actually consumes your GPU
Four things, in decreasing order of predictability:
- Weights. Parameters times bytes per parameter. Fixed and exact.
- KV cache. The formula above. Grows with traffic.
- Activations. Transient working memory during a forward pass. Grows with batch size and prefill chunk size.
- Fragmentation and allocator overhead. Real, awkward, and the reason production frameworks manage memory themselves.
Serving frameworks ask you for a fraction of the card to use, and carve the KV pool out of whatever is left after weights. Set that fraction too high and you get an out-of-memory error during a traffic spike, not at start-up. That is the worst possible time to find out.
Levers that actually work
Grouped-query attention. Not a runtime choice — it is baked into the architecture. It is a model-selection criterion. Check num_key_value_heads before you commit to a model.
KV cache quantisation. Storing the cache in fp8 rather than bf16 halves it. vLLM exposes this as kv_cache_dtype="fp8". Quality impact is small but real, and it should be measured on your own task, not taken from a blog post.
Lower the advertised maximum context. Serving frameworks preallocate against max_model_len. Dropping it from 128k to 8k when your product never needs more is the largest single saving available, and it costs nothing.
Prefix sharing. Identical system prompts across users can point at the same physical blocks. Free, if your prompt layout keeps the shared part first.
More cards. Tensor parallelism splits KV heads across GPUs, so cache per card falls with the parallel degree. It also adds communication cost per token.
Common mistakes
Using parameter count as a proxy for cache size. Shown false above. Read the config.
Forgetting the factor of two for K and V. Half of every home-made estimate on the internet is wrong by exactly this.
Sizing against average context length. The cache is held for the whole request, and your capacity is set by the peak, not the mean. Size against the maximum you allow, then reduce the maximum.
Confusing GiB and GB. A card advertised as 80 GB is 80 GiB of usable HBM in NVIDIA's own reporting. Mixing the two conventions produces a seven-percent error, in exactly the direction that breaks you.
Try it yourself
Add an old multi-head model to the table, using kv_heads = num_attention_heads. Compute how many concurrent 4k sequences it supports on the same 80 GiB card. Then explain the result to somebody as a per-user cost in rupees or dollars per hour of GPU rental.
What to learn next
- PagedAttention — how servers stop wasting most of the memory you calculated above.
- Quantization in practice — the same bytes-per-number lever, applied to weights.
- Latency and throughput — the two numbers this capacity calculation trades between.
Researcher — Mathematics and papers.
The exact expression
For a dense transformer with $L$ layers, $H_{kv}$ key-value heads, head dimension $d_h$, batch $B$, sequence length $S$, and $b$ bytes per stored element:
$$ M_{\text{KV}} = 2 \cdot L \cdot H_{kv} \cdot d_h \cdot S \cdot B \cdot b $$
The leading $2$ counts keys and values. Under multi-head attention $H_{kv} = H$ and $H \cdot d_h = d_{\text{model}}$, which reduces the expression to $2 L d_{\text{model}} S B b$. That is the form most often quoted, and it stops being correct the moment grouped-query attention is used.
Define the grouping ratio $g = H / H_{kv}$. Then
$$ M_{\text{KV}} = \frac{2 L d_{\text{model}} S B b}{g} $$
Llama 3 and Qwen 3 use $g = 4$ to $g = 8$ at the larger sizes.
Cache traffic, which is what actually limits decode
Memory footprint sets capacity. Memory bandwidth sets speed, and these are different constraints.
Each decode step reads the entire cache for the sequences in the batch. Per step:
$$ \text{bytes read} \approx M_{\text{KV}} \quad\text{(the whole cache)} $$
so time per output token has a floor of $M_{\text{KV}} / \text{BW}$. On an H100 at $3.35$ TB/s, a batch holding 40 GiB of cache cannot decode faster than about $12.8$ ms per step. Arithmetic speed is irrelevant to that floor. Attention kernels are memory-bound at decode time for exactly this reason. FlashAttention-style tiling helps by avoiding materialising the score matrix, not by reading fewer cache bytes.
The consequence is a genuine ceiling. Growing batch size raises throughput until cache traffic saturates bandwidth, and beyond that point every extra user slows everyone down proportionally.
Architectural responses
| Approach | Cache per token | Cost |
|---|---|---|
| Multi-head (MHA) | $2 L H d_h b$ | baseline |
| Multi-query (MQA) | $2 L d_h b$ | quality loss at scale |
| Grouped-query (GQA) | $2 L (H/g) d_h b$ | small quality loss, universal in practice |
| Multi-head latent (MLA) | $2 L d_c b$, $d_c \ll H d_h$ | extra arithmetic to reconstruct K and V |
| Sliding window | $2 L H_{kv} d_h w b$, window $w$ | lossy beyond the window |
| Cross-layer sharing | divided by layer group size | quality loss, active research |
MLA (DeepSeek-V2) compresses K and V into a shared latent of dimension $d_c$ and reconstructs per head during attention. Reported cache reduction is on the order of $90\%$ against MHA at comparable quality. The reconstruction can be folded into the query and output projections, so the extra arithmetic is largely free during decode, which is memory-bound anyway.
Cache quantisation has an asymmetric error profile. Keys carry outlier channels that quantise poorly; values are better behaved. Per-channel scaling for keys and per-token scaling for values is the usual compromise (Liu et al., KIVI, 2024, arxiv.org/abs/2402.02750).
Eviction and compression. H2O (Zhang et al., 2023, arxiv.org/abs/2306.14048) keeps a "heavy hitter" subset of tokens, chosen by accumulated attention mass. SnapKV and similar methods compress the prompt's cache after prefill. All of these are lossy, and their reported quality preservation is heavily benchmark-dependent. Treat published retention numbers as an upper bound on what you will see.
Capacity planning, done honestly
$$ B_{\max} = \left\lfloor \frac{M_{\text{GPU}} - M_{\text{weights}} - M_{\text{activations}}}{M_{\text{KV per sequence}}} \right\rfloor $$
Three warnings about applying this.
$M_{\text{KV per sequence}}$ should use the maximum allowed context, not the observed mean, unless your scheduler can preempt. With paged allocation and preemption, planning against a percentile of the length distribution becomes defensible. Preemption itself has a cost, since a preempted sequence is either recomputed or swapped to host memory.
$M_{\text{activations}}$ scales with the prefill chunk size, so it is a scheduler parameter rather than a constant.
And the whole calculation assumes a steady state that real traffic does not have. Size for the peak, then measure preemption rate as your capacity signal.
Papers
- Shazeer, Fast Transformer Decoding, 2019 — arxiv.org/abs/1911.02150
- Ainslie et al., GQA, 2023 — arxiv.org/abs/2305.13245
- Zhang et al., H2O: Heavy-Hitter Oracle for Efficient Generative Inference, 2023 — arxiv.org/abs/2306.14048
- Liu et al., KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache, 2024 — arxiv.org/abs/2402.02750
- DeepSeek-AI, DeepSeek-V2, 2024 — arxiv.org/abs/2405.04434
What to learn next
- PagedAttention — how servers stop wasting most of the memory you calculated above.
- Quantization in practice — the same bytes-per-number lever, applied to weights.
- Latency and throughput — the two numbers this capacity calculation trades between.