Prompt prefix caching
When many requests start with the same long text, a model can process that shared part once and reuse the result, instead of re-reading it for every request.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Prompt prefix caching means a model reads a shared opening section of text once. It reuses that work for every request that starts the same way.
The analogy you have already lived
A tour guide gives the same twenty-minute introduction to every group before splitting off to answer each group's own questions. They do not re-explain the museum's history individually to every single visitor.
They say it once, to the group, then move to what is different — your specific question. A model's shared prompt text (the instructions, the examples, the rules it always follows) is that introduction. It is the same for everyone, so it should not be redone for everyone.
Why it exists
Many real prompts have a prefix — a long opening section of instructions or examples that stays identical across requests. A short part follows that changes: the user's actual question.
A support bot might load 500 words of instructions and examples before ever seeing what the customer typed. If the model reprocesses those 500 words fresh for every customer, most of its work is spent reading the same unchanging text, over and over.
How it works
SHARED instructions (500 words, identical every time)
|
process once, remember the internal "working notes" it produced
|
---------------------------------
| | |
customer 1 customer 2 customer 3
question question question
| | |
reuse the reuse the reuse the
saved notes, saved notes, saved notes,
only process only process only process
THIS question THIS question THIS questionThose "working notes" have a real name: the KV cache. It is internal numbers the model builds while reading text, so it never has to re-read earlier words to understand later ones. Prefix caching means saving that KV cache for the shared part. A copy is handed to every request sharing the same opening text.
A real example you have seen
Chatbots often have a fixed personality or a fixed set of rules: "always be polite, never give medical advice, answer in under 100 words." They pay that setup cost once per unique instruction set. Not once per customer message, if the platform is built well. It is a large part of why some AI products stay affordable at high volume.
The honest part
This only helps when a real prefix is shared word-for-word. Change even one character early in the shared text — a different date at the top, a name filled into a template. Most systems then treat it as a completely different prefix, with no saving at all. Where the variable content sits in your prompt matters more than it seems.
Remember this
- A shared, unchanging prefix can be processed once and reused across many requests.
- The saved work is the model's KV cache for that prefix.
- Put anything that changes after the shared part, not inside it, or the saving disappears.
What to learn next
- The KV cache — the mechanism prefix caching is built on top of.
- Caching embeddings — reusing another kind of repeated computation, further up the pipeline.
- Cutting LLM API costs — where prefix caching fits among the other levers you have.
Developer — Code and libraries.
Setup
pip install transformers torchMeasuring the saving directly
import copy
import time
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("distilgpt2")
model = AutoModelForCausalLM.from_pretrained("distilgpt2")
model.eval()
# A long shared prefix -- stands in for a system prompt every request starts with.
SYSTEM_PROMPT = (
"You are a helpful assistant for a library. Answer briefly and politely. "
"Always mention the due date policy of two weeks. " * 8
)
suffixes = [
"\nUser: Can I renew my book?\nAssistant:",
"\nUser: Do you have any Hindi novels?\nAssistant:",
"\nUser: What time does the library close?\nAssistant:",
]
prefix_ids = tok(SYSTEM_PROMPT, return_tensors="pt").input_ids
print("shared prefix length:", prefix_ids.shape[1], "tokens")
with torch.no_grad():
# Without reuse: every request reprocesses the whole prefix from scratch.
start = time.perf_counter()
for suffix in suffixes:
full_ids = tok(SYSTEM_PROMPT + suffix, return_tensors="pt").input_ids
model(full_ids)
no_reuse_time = time.perf_counter() - start
# With reuse: compute the prefix's KV cache ONCE, replay a copy per request.
start = time.perf_counter()
base_cache = model(prefix_ids, use_cache=True).past_key_values
for suffix in suffixes:
suffix_ids = tok(suffix, return_tensors="pt").input_ids
model(suffix_ids, past_key_values=copy.deepcopy(base_cache), use_cache=True)
reuse_time = time.perf_counter() - start
print(f"without prefix cache: {no_reuse_time*1000:7.1f} ms for {len(suffixes)} requests")
print(f"with prefix cache: {reuse_time*1000:7.1f} ms for {len(suffixes)} requests")
print(f"speedup: {no_reuse_time/reuse_time:.2f}x")shared prefix length: 193 tokens without prefix cache: 817.5 ms for 3 requests with prefix cache: 223.6 ms for 3 requests speedup: 3.66x
This ran on CPU, on this one machine, with a small model and a short prefix. Real numbers, not projections. A larger model, a longer shared prefix, or a GPU changes the exact figures. The direction holds generally: reusing the prefix is faster, and more so the longer the prefix is. It is skipping real, measured work.
Line-by-line walkthrough
use_cache=True. Tells the model to return its internal KV cache alongside the normal output, instead of discarding it.
copy.deepcopy(base_cache). Each request gets its own independent copy of the cached prefix state. Reusing the same object across requests without copying can let one request's continuation corrupt another's — an easy, quiet bug.
Why the gap grows with prefix length. The saved work is roughly proportional to the number of prefix tokens. A one-sentence prefix barely matters. A 2,000-token system prompt with examples can dominate total cost if recomputed every time.
Common mistakes
Putting variable content inside the prefix. A date, a username, or a session ID typed into the middle of the "shared" text makes every request's prefix unique. The cache never reuses anything. Keep the truly shared part first, and put anything that varies after it.
Assuming this is free memory-wise. A cached KV state takes real memory, proportional to prefix length and model size. See how much memory the KV cache eats. Caching prefixes for many different system prompts at once can use more memory than expected.
Building this by hand when a server already does it. Serving frameworks like vLLM and NVIDIA's TGI/Triton implement automatic prefix caching internally. They match shared prefixes across arbitrary requests, with no logic for you to write. Reach for those before hand-rolling it in production.
Forgetting this is not the same as the semantic cache. Semantic caching skips the model entirely for a repeated meaning. Prefix caching still runs the model on every request — it only skips redoing the shared opening text. They solve different problems and combine well together.
Try it yourself
Double SYSTEM_PROMPT's repeat count from 8 to 16 and re-run. Watch how the speedup changes as the shared prefix gets longer relative to the short, unique suffix.
What to learn next
- The KV cache — the mechanism prefix caching is built on top of.
- Caching embeddings — reusing another kind of repeated computation, further up the pipeline.
- Cutting LLM API costs — where prefix caching fits among the other levers you have.
Researcher — Mathematics and papers.
What is actually being reused
For a transformer with $L$ layers and $H$ attention heads, processing a sequence of length $n$ produces, for every layer and head, a key vector $k_i$ and value vector $v_i$ for each token $i$. Autoregressive generation needs $k_i, v_i$ for every earlier token to compute attention for the current one, so caching them avoids recomputing the full $n \times n$ attention pattern from scratch at each new token — this is the KV cache mechanism covered in the KV cache.
Prefix caching is the observation that if a prefix of length $p$ is byte-identical across requests, its KV tensors are also identical across requests (given a deterministic forward pass), and can be computed once and copied or referenced rather than recomputed. The compute saved per request is approximately
$$\Delta \approx c \cdot p$$
where $c$ is the per-token forward-pass cost. Because attention cost scales with sequence length, a longer shared prefix produces a larger absolute saving, which is why system prompts with long few-shot examples benefit the most.
Radix-tree prefix sharing
Naive prefix caching matches one exact prefix. vLLM's automatic prefix caching (Kwon et al., 2023, and subsequent releases) generalises this with a radix tree over token sequences. Any two requests sharing any common prefix — not just a single hardcoded system prompt — automatically share the cached portion, down to the token where they diverge. This turns prefix caching from a manual optimisation into a property of the serving system.
SGLang's RadixAttention (Zheng et al., 2024) takes the same structure further, sharing cache across a full tree of branching generations. This helps tree-of-thought and multi-branch agent workloads, where many continuations share a long common ancestor.
Memory cost
Cached KV state consumes memory proportional to $L \times H \times p \times d_{\text{head}} \times 2$ (the factor of 2 for keys and values), per cached sequence. Serving systems bound this with an eviction policy over the radix tree, typically LRU over tree nodes. Block-based memory management (PagedAttention; Kwon et al., 2023) reclaims memory at block granularity, rather than requiring one large contiguous allocation per cached prefix.
Papers and systems
- Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM), SOSP 2023 — arxiv.org/abs/2309.06180
- Zheng et al., SGLang: Efficient Execution of Structured Language Model Programs, 2024 — arxiv.org/abs/2312.07104
- Gim et al., Prompt Cache: Modular Attention Reuse for Low-Latency Inference, 2023 — arxiv.org/abs/2311.04934
What to learn next
- The KV cache — the mechanism prefix caching is built on top of.
- Caching embeddings — reusing another kind of repeated computation, further up the pipeline.
- Cutting LLM API costs — where prefix caching fits among the other levers you have.