Caching and Cost Control

Prompt prefix caching

When many requests start with the same long text, a model can process that shared part once and reuse the result, instead of re-reading it for every request.

Read these first

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Prompt prefix caching means a model reads a shared opening section of text once. It reuses that work for every request that starts the same way.

The analogy you have already lived

A tour guide gives the same twenty-minute introduction to every group before splitting off to answer each group's own questions. They do not re-explain the museum's history individually to every single visitor.

They say it once, to the group, then move to what is different — your specific question. A model's shared prompt text (the instructions, the examples, the rules it always follows) is that introduction. It is the same for everyone, so it should not be redone for everyone.

Why it exists

Many real prompts have a prefix — a long opening section of instructions or examples that stays identical across requests. A short part follows that changes: the user's actual question.

A support bot might load 500 words of instructions and examples before ever seeing what the customer typed. If the model reprocesses those 500 words fresh for every customer, most of its work is spent reading the same unchanging text, over and over.

How it works

   SHARED instructions (500 words, identical every time)
              |
   process once, remember the internal "working notes" it produced
              |
       ---------------------------------
       |              |               |
   customer 1      customer 2      customer 3
   question        question        question
       |              |               |
   reuse the       reuse the       reuse the
   saved notes,    saved notes,    saved notes,
   only process    only process    only process
   THIS question   THIS question   THIS question

Those "working notes" have a real name: the KV cache. It is internal numbers the model builds while reading text, so it never has to re-read earlier words to understand later ones. Prefix caching means saving that KV cache for the shared part. A copy is handed to every request sharing the same opening text.

A real example you have seen

Chatbots often have a fixed personality or a fixed set of rules: "always be polite, never give medical advice, answer in under 100 words." They pay that setup cost once per unique instruction set. Not once per customer message, if the platform is built well. It is a large part of why some AI products stay affordable at high volume.

The honest part

This only helps when a real prefix is shared word-for-word. Change even one character early in the shared text — a different date at the top, a name filled into a template. Most systems then treat it as a completely different prefix, with no saving at all. Where the variable content sits in your prompt matters more than it seems.

Remember this

  • A shared, unchanging prefix can be processed once and reused across many requests.
  • The saved work is the model's KV cache for that prefix.
  • Put anything that changes after the shared part, not inside it, or the saving disappears.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install transformers torch

Measuring the saving directly

prefix_cache.py
import copy
import time
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("distilgpt2")
model = AutoModelForCausalLM.from_pretrained("distilgpt2")
model.eval()

# A long shared prefix -- stands in for a system prompt every request starts with.
SYSTEM_PROMPT = (
    "You are a helpful assistant for a library. Answer briefly and politely. "
    "Always mention the due date policy of two weeks. " * 8
)
suffixes = [
    "\nUser: Can I renew my book?\nAssistant:",
    "\nUser: Do you have any Hindi novels?\nAssistant:",
    "\nUser: What time does the library close?\nAssistant:",
]
prefix_ids = tok(SYSTEM_PROMPT, return_tensors="pt").input_ids
print("shared prefix length:", prefix_ids.shape[1], "tokens")

with torch.no_grad():
    # Without reuse: every request reprocesses the whole prefix from scratch.
    start = time.perf_counter()
    for suffix in suffixes:
        full_ids = tok(SYSTEM_PROMPT + suffix, return_tensors="pt").input_ids
        model(full_ids)
    no_reuse_time = time.perf_counter() - start

    # With reuse: compute the prefix's KV cache ONCE, replay a copy per request.
    start = time.perf_counter()
    base_cache = model(prefix_ids, use_cache=True).past_key_values
    for suffix in suffixes:
        suffix_ids = tok(suffix, return_tensors="pt").input_ids
        model(suffix_ids, past_key_values=copy.deepcopy(base_cache), use_cache=True)
    reuse_time = time.perf_counter() - start

print(f"without prefix cache: {no_reuse_time*1000:7.1f} ms for {len(suffixes)} requests")
print(f"with prefix cache:    {reuse_time*1000:7.1f} ms for {len(suffixes)} requests")
print(f"speedup: {no_reuse_time/reuse_time:.2f}x")
Output
shared prefix length: 193 tokens
without prefix cache:   817.5 ms for 3 requests
with prefix cache:      223.6 ms for 3 requests
speedup: 3.66x

This ran on CPU, on this one machine, with a small model and a short prefix. Real numbers, not projections. A larger model, a longer shared prefix, or a GPU changes the exact figures. The direction holds generally: reusing the prefix is faster, and more so the longer the prefix is. It is skipping real, measured work.

Line-by-line walkthrough

use_cache=True. Tells the model to return its internal KV cache alongside the normal output, instead of discarding it.

copy.deepcopy(base_cache). Each request gets its own independent copy of the cached prefix state. Reusing the same object across requests without copying can let one request's continuation corrupt another's — an easy, quiet bug.

Why the gap grows with prefix length. The saved work is roughly proportional to the number of prefix tokens. A one-sentence prefix barely matters. A 2,000-token system prompt with examples can dominate total cost if recomputed every time.

Common mistakes

Putting variable content inside the prefix. A date, a username, or a session ID typed into the middle of the "shared" text makes every request's prefix unique. The cache never reuses anything. Keep the truly shared part first, and put anything that varies after it.

Assuming this is free memory-wise. A cached KV state takes real memory, proportional to prefix length and model size. See how much memory the KV cache eats. Caching prefixes for many different system prompts at once can use more memory than expected.

Building this by hand when a server already does it. Serving frameworks like vLLM and NVIDIA's TGI/Triton implement automatic prefix caching internally. They match shared prefixes across arbitrary requests, with no logic for you to write. Reach for those before hand-rolling it in production.

Forgetting this is not the same as the semantic cache. Semantic caching skips the model entirely for a repeated meaning. Prefix caching still runs the model on every request — it only skips redoing the shared opening text. They solve different problems and combine well together.

Try it yourself

Double SYSTEM_PROMPT's repeat count from 8 to 16 and re-run. Watch how the speedup changes as the shared prefix gets longer relative to the short, unique suffix.

What to learn next

Researcher — Mathematics and papers.

What is actually being reused

For a transformer with $L$ layers and $H$ attention heads, processing a sequence of length $n$ produces, for every layer and head, a key vector $k_i$ and value vector $v_i$ for each token $i$. Autoregressive generation needs $k_i, v_i$ for every earlier token to compute attention for the current one, so caching them avoids recomputing the full $n \times n$ attention pattern from scratch at each new token — this is the KV cache mechanism covered in the KV cache.

Prefix caching is the observation that if a prefix of length $p$ is byte-identical across requests, its KV tensors are also identical across requests (given a deterministic forward pass), and can be computed once and copied or referenced rather than recomputed. The compute saved per request is approximately

$$\Delta \approx c \cdot p$$

where $c$ is the per-token forward-pass cost. Because attention cost scales with sequence length, a longer shared prefix produces a larger absolute saving, which is why system prompts with long few-shot examples benefit the most.

Radix-tree prefix sharing

Naive prefix caching matches one exact prefix. vLLM's automatic prefix caching (Kwon et al., 2023, and subsequent releases) generalises this with a radix tree over token sequences. Any two requests sharing any common prefix — not just a single hardcoded system prompt — automatically share the cached portion, down to the token where they diverge. This turns prefix caching from a manual optimisation into a property of the serving system.

SGLang's RadixAttention (Zheng et al., 2024) takes the same structure further, sharing cache across a full tree of branching generations. This helps tree-of-thought and multi-branch agent workloads, where many continuations share a long common ancestor.

Memory cost

Cached KV state consumes memory proportional to $L \times H \times p \times d_{\text{head}} \times 2$ (the factor of 2 for keys and values), per cached sequence. Serving systems bound this with an eviction policy over the radix tree, typically LRU over tree nodes. Block-based memory management (PagedAttention; Kwon et al., 2023) reclaims memory at block granularity, rather than requiring one large contiguous allocation per cached prefix.

Papers and systems

What to learn next