Cache invalidation for RAG
When the documents behind a RAG answer change, every cached answer built from the old document has to be found and thrown out, not left to quietly go stale.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
When a source document changes, every cached answer built from its old version has to be thrown out. It gets rebuilt.
The analogy you have already lived
A school notice board says the exam is on Monday. A student photographs it and tells five friends. The school then moves the exam to Wednesday and updates the board.
If nobody tells the five friends, they show up Monday with an outdated screenshot that used to be correct. The board was updated. Their copy of the old information was not.
RAG systems have exactly this problem, at a much larger scale.
Why it exists
RAG stands for retrieval-augmented generation. A model is given relevant documents alongside a question, and answers using those documents instead of only its own memory. What is RAG? covers this in full.
Earlier lessons cached answers to save time and money. RAG answers are built from real documents — a policy page, a product manual, a support article. Those documents change. A cached answer built from yesterday's document is now built from a lie. Nothing about the caching system itself broke.
How it works
Monday: document says "refunds take 5-7 days"
|
a question gets answered using that document
|
the answer gets CACHED, tied to that version of the document
|
-----------------------------------------------------
Wednesday: someone edits the document
|
"refunds take 5-7 days" changes to "refunds take 2-3 days"
|
every cached answer built from the OLD text must be thrown out
|
the next question gets answered fresh, from the NEW documentThe trick is tying every cached answer to which version of the document produced it. Not only to the question that was asked.
A real example you have seen
A company chatbot that confidently tells a customer the old return policy, days after the policy page was updated. This is not a rare bug. It is one of the most common failure modes in real RAG deployments, and it happens exactly when caching is not tied to document freshness.
The honest part
There is no setting that makes this problem disappear. Either you check freshness on every single answer, which removes most of the caching benefit. Or you accept some window of staleness, in exchange for speed and cost — and choose that window on purpose.
Remember this
- A cached RAG answer is only correct as long as the document it came from has not changed.
- Tie every cached answer to a specific version of its source documents, not just the question.
- Some staleness window is often unavoidable. Choose it deliberately, and know what it is.
What to learn next
- What is RAG? — the retrieval pipeline whose freshness this lesson protects.
- Document ingestion pipelines for RAG — where documents enter the system in the first place.
- Semantic caching — the caching layer this lesson adds version-awareness to.
Developer — Code and libraries.
Setup
No installs needed — this uses only Python's standard library hashlib.
Tying cache entries to document versions
import hashlib
documents = {
"policy.md": "Refunds take 5-7 business days.",
"hours.md": "Support is open 9am to 6pm on weekdays.",
}
def doc_hash(text: str) -> str:
return hashlib.sha256(text.encode()).hexdigest()[:8]
answer_cache = {} # (question, doc_id, doc_hash) -> answer
depends_on = {} # doc_id -> set of cache keys built from it
def fake_llm_answer(question, doc_text):
return f"Based on the document: {doc_text}"
def answer_question(question, doc_id):
text = documents[doc_id]
key = (question, doc_id, doc_hash(text)) # the hash makes an edit a new key
if key in answer_cache:
return answer_cache[key], "HIT"
answer = fake_llm_answer(question, text)
answer_cache[key] = answer
depends_on.setdefault(doc_id, set()).add(key)
return answer, "MISS"
def update_document(doc_id, new_text):
documents[doc_id] = new_text
stale_keys = depends_on.pop(doc_id, set())
for key in stale_keys:
answer_cache.pop(key, None) # purge every answer built from the old text
return len(stale_keys)
q = "How long do refunds take?"
answer, status = answer_question(q, "policy.md")
print(f"[{status}] {answer}")
answer, status = answer_question(q, "policy.md")
print(f"[{status}] {answer}")
print("cache size before update:", len(answer_cache))
purged = update_document("policy.md", "Refunds take 2-3 business days after a policy change.")
print(f"document updated, purged {purged} stale cache entry")
print("cache size after update:", len(answer_cache))
answer, status = answer_question(q, "policy.md")
print(f"[{status}] {answer}")[MISS] Based on the document: Refunds take 5-7 business days. [HIT] Based on the document: Refunds take 5-7 business days. cache size before update: 1 document updated, purged 1 stale cache entry cache size after update: 0 [MISS] Based on the document: Refunds take 2-3 business days after a policy change.
The same question, asked before and after the document edit, correctly gets two different answers. The second one is a fresh MISS, not a stale HIT.
Line-by-line walkthrough
The cache key includes doc_hash(text). This is the entire mechanism. When the document's text changes, its hash changes, so the old key can never match again. The stale answer is not deleted immediately by this line alone. It becomes unreachable — which is why the depends_on index exists too.
depends_on. A reverse index: for each document, which cache keys were built from it. Without this, an edited document's old cache entries would sit in memory forever, unreachable but never cleaned up. A slow memory leak, instead of a correctness bug.
update_document purges by document, not by guessing which questions are affected. This is deliberately conservative. Every cached answer that touched the changed document is thrown out, even if the specific edit would not have changed that particular answer. Being wrong in the direction of "recompute a bit too much" is far safer. Being wrong in the direction of "serve stale information" is not.
Common mistakes
Caching by question text alone. If the key is only question, an edited document is invisible to the cache entirely. The stale answer is served forever, correctly matching the exact question that was asked.
A time-based TTL as the only defence. A TTL (time-to-live) helps, but a document can change five minutes after being cached, and one minute before the TTL expires. That leaves up to the full TTL window of staleness, with no way to shorten it. Version-based invalidation, as above, catches the change immediately, whenever it happens.
Invalidating only the exact chunk that changed. In a real RAG system, a document is usually split into chunks for retrieval. An edit to one paragraph can change which chunks get retrieved for a given question, not only the content of the edited chunk. So invalidating "only the changed chunk" can still miss answers assembled from a now-different retrieval set.
No invalidation path for deletions. A document being removed entirely needs the same purge as an edit. Forgetting this leaves cached answers built from content that no longer exists anywhere.
Try it yourself
Add a second document, and a question whose answer combines both documents. Update only one of the two. Check whether your cache correctly purges the combined answer, not just answers about the single document that changed.
What to learn next
- What is RAG? — the retrieval pipeline whose freshness this lesson protects.
- Document ingestion pipelines for RAG — where documents enter the system in the first place.
- Semantic caching — the caching layer this lesson adds version-awareness to.
Researcher — Mathematics and papers.
Freshness as a correctness property, not a performance one
Treat cache staleness in RAG as a correctness bound, not a performance detail. The maximum time between a document changing and every dependent cached answer being invalidated is a number you can measure and should report — the same way you would report a latency SLO. Push-based invalidation — the document store notifies the cache on every write — gives a bound close to zero. Poll-based invalidation, re-checking document hashes periodically, gives a bound equal to the poll interval. TTL-only invalidation gives a bound equal to the TTL, regardless of when the actual edit happened.
Dependency tracking at scale
The depends_on reverse index in the demo is a simplified version of what database systems call materialized view maintenance. A materialized view — the cached answer — is a precomputed function of underlying base tables, the documents. The system needs to know exactly which views depend on which rows, to invalidate correctly and minimally on a write. The general problem — computing the minimal set of dependents to invalidate — is well studied in that literature (Gupta and Mumick, 1995). The same principles transfer directly: track dependencies explicitly, invalidate by dependency rather than by blanket expiry, and prefer over-invalidating to under-invalidating when the dependency graph is uncertain.
Chunk-level versus document-level invalidation
For a chunked RAG index, invalidation can operate at three granularities, each a different trade-off:
- Document-level — any change to the source document invalidates everything derived from it. Simple, safe, coarse; can cause unnecessary recomputation for large documents with small edits.
- Chunk-level — only cached answers that retrieved the specific changed chunk are invalidated. Finer-grained, but it requires tracking which chunks each cached answer actually retrieved. It can still miss answers where the set of retrieved chunks changes, because a new chunk now ranks higher — even though no chunk it previously retrieved changed at all.
- Retrieval-set-level — invalidate only when the top-$k$ retrieval results for a cached question's embedding would actually change. Most precise, most expensive to check, since it requires rerunning retrieval to find out.
Most production systems default to document-level or chunk-level invalidation and accept the residual risk of the retrieval-set edge case, monitored rather than eliminated.
Reading
- Gupta and Mumick, Maintenance of Materialized Views: Problems, Techniques, and Applications, IEEE Data Engineering Bulletin, 1995
What to learn next
- What is RAG? — the retrieval pipeline whose freshness this lesson protects.
- Document ingestion pipelines for RAG — where documents enter the system in the first place.
- Semantic caching — the caching layer this lesson adds version-awareness to.