Prompt caching
In one sentence Prompt caching stores the computed state of a prompt's repeated prefix, so identical opening tokens are billed and processed at a fraction of the cost.
Updated
Prompt caching reuses the model's internal processing of a prompt prefix it has seen before, cutting both the cost and the latency of repeated content.
A security guard checks your full ID on day one. By Thursday, he waves you through — the verification work is done and remembered, and only whatever is new (today's visitor) needs processing. LLM serving has the same opportunity: most applications resend an identical opening on every call — the same system-prompt, the same tool definitions, the same reference document — followed by a short new question.
The mechanism underneath is the KV cache: processing a prompt fills that cache token by token, and identical prefix tokens always produce identical cache contents. So the provider stores the cache for your prefix and, on the next request, resumes from where it ends rather than recomputing from token one.
The economics are dramatic enough to design around. Cached input tokens are billed at roughly a tenth of the normal price (varies by provider — Anthropic ~10×, OpenAI ~2×, Gemini similar in spirit), and time-to-first-token drops sharply for long prefixes. The one design rule that unlocks it: caching matches exact prefixes, so put stable content first and variable content last.
[system prompt][tool schemas][big reference doc] ← stable: cached
[conversation so far][new user message] ← variable: comes afterA timestamp or user name placed at the top of the system prompt silently breaks the match, and with it your savings. Caches expire within minutes to hours; some APIs need explicit cache-control markers, others cache automatically.
Where to go next
- Full lesson: vLLM
- Related terms: kv-cache, context-window, system-prompt, latency