Giving each chunk its context back
A chunk pulled out of a document often loses the context that made it meaningful, and prepending a short header naming its source fixes both retrieval and readability.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A contextual chunk header is a short label added to a chunk, naming the document and section it came from.
Think about a photocopied page torn from the middle of a book. No title, no chapter name, nothing at the top. You cannot tell which book it is from, unless someone writes that at the top for you.
Why it exists
A chunk like "It grew by 6 percent that quarter, driven by strong demand" reads fine inside its document. Pulled out alone, "it" means nothing. Which company? Which quarter? An earlier sentence answered those questions, and it did not make it into this chunk.
This hurts in two places at once. A human reading the chunk cannot tell what it is about. A search system embedding the chunk also struggles. Its embedding is built purely from the chunk's own words, and those words are missing the topic.
The fix is direct. Write the missing context back in, at the top of the chunk, before anything else happens to it.
How it works
Bare chunk:
"It grew by 6 percent that quarter, driven by
strong demand in North America."
With a context header added:
"Document: Infosys Q3 2023 earnings report. Section: Revenue.
It grew by 6 percent that quarter, driven by
strong demand in North America."The header turns a floating, unlabelled sentence into something that stands on its own. That helps a person reading it, and a search system trying to match it to a question.
Where you have already seen it
- Email previews. Gmail shows the sender and subject above the message snippet, because the snippet alone is often unclear.
- News aggregator apps. A quoted paragraph usually comes with the source name and headline attached above it.
- Search engine results. Google shows the page title and site name above every snippet, not the snippet alone.
Remember this
- A chunk pulled out of context often loses the words that made it meaningful.
- Prepending a short header — document title, section name — restores that missing context.
- This helps both the search step and the person reading the final answer.
What to learn next
- Chunking by document structure — where a breadcrumb, one kind of header, comes from.
- Semantic chunking — a companion technique this pairs well with.
- Retrieve small, return big — a different fix for the same underlying problem.
Developer — Code and libraries.
This compares the embedding similarity of a bare chunk against a question, with and without a prepended context header.
Setup
pip install transformers torchThe first run downloads sentence-transformers/all-MiniLM-L6-v2, about 90 MB.
Measuring the effect of a context header
import torch
from transformers import AutoTokenizer, AutoModel
tok = AutoTokenizer.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model = AutoModel.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model.eval()
def embed(texts):
enc = tok(texts, padding=True, truncation=True, return_tensors="pt")
with torch.no_grad():
out = model(**enc)
mask = enc["attention_mask"].unsqueeze(-1).float()
pooled = (out.last_hidden_state * mask).sum(1) / mask.sum(1)
return torch.nn.functional.normalize(pooled, dim=1)
query = "How much did Infosys revenue grow in the third quarter of 2023?"
chunk_bare = "It grew by 6 percent that quarter, driven by strong demand in North America."
chunk_with_header = (
"Document: Infosys Q3 2023 earnings report. Section: Revenue. "
"It grew by 6 percent that quarter, driven by strong demand in North America."
)
vecs = embed([query, chunk_bare, chunk_with_header])
sim_bare = (vecs[0] * vecs[1]).sum().item()
sim_header = (vecs[0] * vecs[2]).sum().item()
print(f"similarity, bare chunk: {sim_bare:.3f}")
print(f"similarity, with context header: {sim_header:.3f}")similarity, bare chunk: 0.428 similarity, with context header: 0.783
Model weights are pinned, so these numbers are exact for this model version. A different sentence-transformers/all-MiniLM-L6-v2 revision could shift the third decimal.
Line by line
embed() reuses the mean-pooling pattern from the semantic chunking lesson — one vector per input text, normalised to length 1, so a dot product gives cosine similarity directly.
The similarity jump, 0.428 to 0.783, is the entire result. Nothing about the chunk's actual claim changed — "it grew by 6 percent" is still exactly what it says. What changed is whether the embedding model has any signal connecting this text to "Infosys" and "revenue," the words the query actually uses.
This is a retrieval fix, not a generation fix. The header helps the chunk get found in the first place. Once found, a language model reading the chunk still benefits from the same header, for the same reason a human would.
Common mistakes
Writing a header so long it dominates the chunk. A three-paragraph header on a two-sentence chunk skews the embedding toward the header's wording, not the chunk's actual content. Keep headers to one short line.
Generating headers with an expensive model call per chunk, at query time. Headers are a property of the document, not the query. Generate them once, when the document is chunked, and store the result — not on every search request.
Using a header that duplicates information already obvious in the chunk. A chunk that already says "Infosys reported Q3 2023 revenue of..." gains little from a header repeating "Document: Infosys Q3 2023 report." Reserve headers for context the chunk itself does not restate.
Assuming this fully replaces good chunk boundaries. A context header repairs a chunk that lost its topic. It does not repair a chunk that was cut mid-sentence. Combine this with the chunking methods from earlier lessons, not instead of them.
Try it yourself
Change chunk_with_header to use a wrong company name — "Document: Wipro Q3 2023 earnings report" — and re-run.
Similarity to the Infosys-related query should drop compared to the correct header, even though the sentence about "6 percent growth" is unchanged. The header actively steers the embedding. A wrong header actively steers it wrong, which is worth knowing before automating header generation from possibly-unreliable metadata.
What to learn next
- Semantic chunking — building the chunks a header gets added to.
- Sentence transformers — more on the embedding model used here.
- Bi-encoders vs cross-encoders — how this kind of embedding gets used in a full search system.
Researcher — Mathematics and papers.
Framing as context injection
A context header is a case of context injection: adding text to a unit before embedding, where that text is true of the unit but not literally present in its original span. Structure-aware breadcrumbs, covered earlier in this section, are one source of injected context. Document metadata — title, author, date, source system — is another. The technique generalises beyond headings to anything the surrounding document or its metadata can supply.
Anthropic's contextual retrieval
Anthropic (2024), Introducing Contextual Retrieval (company engineering blog, not a peer-reviewed paper), popularised a specific version of this idea. A language model generates a short, chunk-specific explanatory sentence — beyond a static header, a summary of how this particular chunk relates to the whole document — prepended before embedding and before building a lexical (BM25) index. Their reported results showed retrieval failure rates dropping substantially when this was combined with reranking, on their internal evaluation sets. Treat these as one team's reported numbers on their own benchmarks, not an independently replicated, peer-reviewed result.
Why this works: an information-theoretic framing
A bi-encoder embedding compresses a fixed-length passage into a fixed-size vector, computed from that passage's own tokens. Zero information the model was not given can appear in the embedding. A chunk whose original document context supplied a coreference ("it" meaning "Infosys") but which no longer contains that resolution has, in a real sense, lost that information from the embedding's input entirely. Prepending it restores an input signal the embedding model needs. This is not a trick or a workaround; it directly addresses the bi-encoder's core constraint of compressing only what it is shown. See Bi-encoders vs cross-encoders for how this constraint compares to a cross-encoder's very different, and more expensive, access to full context at scoring time.
Cost trade-off
Static headers (breadcrumb, title, section) cost nothing beyond string concatenation. LLM-generated contextual summaries, as in Anthropic's approach, cost one model call per chunk at indexing time — a one-time cost proportional to corpus size, not to query volume, since it happens during chunking rather than at search time. That sets it apart from query-time techniques like query expansion, whose cost recurs on every search.
Key references
- Anthropic (2024). Introducing Contextual Retrieval. Anthropic engineering blog — a practitioner report, not a peer-reviewed publication.
- Gao, Y. et al. (2023). Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997
Current state and open problems
Whether to hand-write header templates, generate them with a cheap rule, or spend a full LLM call per chunk is an unresolved cost-quality trade-off, decided per project rather than by a settled best practice. There is no widely reproduced, controlled comparison isolating the contribution of context headers alone, independent of the reranking and hybrid-search components frequently deployed alongside them in reported case studies. Attributing a quality improvement to the header specifically, versus the rest of the pipeline it was tested inside, remains a genuine methodological gap in the public literature on this technique.
What to learn next
- Structure-aware chunking — one concrete source of header content.
- Hybrid search — the companion technique context headers are frequently combined with.
- Building a retrieval eval set — measuring whether a header change actually helped.