Chunking and Long Documents

Semantic chunking

Semantic chunking splits a document where its topic actually shifts, using meaning rather than a fixed character count to decide where one chunk ends and the next begins.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Semantic chunking splits a document where the topic changes, not at a fixed character count.

Think about sorting a mixed box of family photographs into piles. Birthday photos go in one pile, a wedding in another, a holiday trip in a third. You are not counting photos into equal stacks. You are grouping them by what each one is actually about.

Why it exists

The chunking methods from the last lesson cut by character count or by sentence boundary. Both ignore meaning entirely.

That causes a real problem. Two unrelated sentences can land in the same chunk, only because they both fit under the size limit. A search system then retrieves that chunk for a question about one topic. It drags in irrelevant text from the other.

Semantic chunking fixes this by measuring how similar each sentence is to the one before it. A large drop in similarity signals a topic change — a good place to cut.

How it works

Each sentence gets turned into a list of numbers, called an embedding, that captures its meaning. Sentences about similar things get similar numbers.

Sentence:     [about interest rates]  [about interest rates]  [about monsoon rainfall]
Similarity
to next:              0.83                    0.15  <- big drop, cut here
                                                0.45

Result:
  Chunk A: the two interest-rate sentences, kept together
  Chunk B: the rainfall sentence, starts a new chunk

A small drop in similarity between two sentences means they likely belong together. A big drop means the topic moved on, and that is where the chunk boundary goes.

Where you have already seen it

  • Podcast and video apps with auto-chapters. A single episode gets split into named sections, often based on where the topic in the transcript shifts.
  • Meeting-notes tools. Notes from an hour-long call get grouped under separate headings per discussion topic, not by the clock.
  • Long WhatsApp or email threads, summarised. A summary tool that separates "budget discussion" from "hiring update" within one thread is doing something similar.

Remember this

  • Semantic chunking measures meaning, using embeddings, rather than counting characters alone.
  • A big drop in sentence-to-sentence similarity marks a good place to cut.
  • It costs more compute than fixed-size chunking, since every sentence needs an embedding first.

What to learn next

Developer — Code and libraries.

This builds real sentence embeddings with a small pretrained model, then finds the breakpoint where topic similarity drops.

Setup

bash
pip install transformers torch

The first run downloads sentence-transformers/all-MiniLM-L6-v2, about 90 MB. It caches locally after that.

Finding the semantic breakpoint

semantic_chunk.py
import re
import torch
from transformers import AutoTokenizer, AutoModel

doc = (
    "The Reserve Bank of India sets the repo rate. "
    "The repo rate is the interest rate at which the RBI lends to commercial banks. "
    "A higher repo rate makes loans more expensive for everyone. "
    "Monsoon rainfall this year was above average across most states. "
    "Good rainfall usually means a strong harvest for rice and wheat farmers. "
    "A strong harvest tends to keep food prices low in local markets."
)
sentences = re.split(r'(?<=[.!?])\s+', doc.strip())

tok = AutoTokenizer.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model = AutoModel.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model.eval()

enc = tok(sentences, padding=True, truncation=True, return_tensors="pt")
with torch.no_grad():
    out = model(**enc)

# mean-pool the token embeddings into one vector per sentence
mask = enc["attention_mask"].unsqueeze(-1).float()
pooled = (out.last_hidden_state * mask).sum(1) / mask.sum(1)
pooled = torch.nn.functional.normalize(pooled, dim=1)

sims = (pooled[:-1] * pooled[1:]).sum(dim=1)   # cosine similarity, consecutive pairs

for i, s in enumerate(sentences):
    tail = f"   similarity to next: {sims[i]:.3f}" if i < len(sims) else ""
    print(f"{i}: {s}{tail}")

threshold = sims.mean() - sims.std()
breakpoints = [i for i, s in enumerate(sims) if s < threshold]
print(f"\nbreakpoint threshold: {threshold:.3f}")
print(f"split after sentence index: {breakpoints}")
Output
0: The Reserve Bank of India sets the repo rate.   similarity to next: 0.829
1: The repo rate is the interest rate at which the RBI lends to commercial banks.   similarity to next: 0.635
2: A higher repo rate makes loans more expensive for everyone.   similarity to next: 0.149
3: Monsoon rainfall this year was above average across most states.   similarity to next: 0.398
4: Good rainfall usually means a strong harvest for rice and wheat farmers.   similarity to next: 0.448
5: A strong harvest tends to keep food prices low in local markets.

breakpoint threshold: 0.236
split after sentence index: [2]

Model weights are pinned, so the numbers above are exact for this model version — a different transformers release could shift them slightly at the third decimal.

Line by line

pooled = (out.last_hidden_state * mask).sum(1) / mask.sum(1) builds one vector per sentence by averaging its token vectors, ignoring padding. This is mean pooling, the same operation the sentence-transformers library performs internally for this model family.

normalize(pooled, dim=1) scales each vector to length 1. After that, a dot product between two vectors equals their cosine similarity, which is why the next line can skip a separate cosine-similarity function.

sims = (pooled[:-1] * pooled[1:]).sum(dim=1) multiplies every sentence's vector against the next sentence's vector, element-wise, then sums. That is a dot product computed for every consecutive pair at once, with no loop.

The threshold, mean - std, is one common rule of thumb: cut wherever similarity drops more than one standard deviation below the document's average. It found the real topic boundary here — between "makes loans more expensive" and "monsoon rainfall" — with no keyword rule telling it where to look.

Common mistakes

Running this on a two- or three-sentence document. Standard deviation over three numbers is close to meaningless, and the threshold becomes unstable. Semantic chunking needs a reasonable number of sentences to have a stable notion of "average similarity."

Assuming the threshold formula is universal. mean - std is one convention among several. Some implementations use a percentile cutoff instead, such as the 5th percentile of all pairwise similarities. Test both on your own documents.

Re-embedding every sentence for every new document, live, in a user-facing request. Embedding is cheap per sentence but adds up. Production systems chunk documents once, offline, and store the result — not on every search.

Forgetting this still needs a maximum chunk size. A long run of highly similar sentences never triggers a break on its own. Pair semantic chunking with a hard size cap, so one topic that runs for ten pages does not become one giant chunk.

Try it yourself

Add a fourth topic to doc — one or two more sentences about something else entirely, like a cricket score. Re-run the script.

You should see a new similarity drop appear right before your added sentences, and a new entry in breakpoints. The threshold is recomputed from the whole document each time, so adding text changes the mean and standard deviation slightly, not only where you added the sentence.

What to learn next

Researcher — Mathematics and papers.

Formal definition

Given a sequence of sentence embeddings e_1, ..., e_n, define the adjacent similarity series:

text
sim_i = cos(e_i, e_{i+1}) = (e_i . e_{i+1}) / (||e_i|| * ||e_{i+1}||),   i = 1..n-1
  • e_i is the pooled embedding of sentence i.
  • cos(.,.) is cosine similarity, in [-1, 1] for arbitrary vectors and typically [0, 1] for sentence embeddings from models trained with a contrastive objective.

A breakpoint is any index i where sim_i falls below a threshold tau. The threshold is a design choice, not derived from theory. Common conventions:

  • Standard-deviation rule: tau = mean(sim) - k * std(sim), with k around 1.
  • Percentile rule: tau set to the p-th percentile of the sim distribution, p typically between 5 and 25.
  • Local-minimum rule: cut only at sentences that are a local minimum in sim, which avoids over-splitting a document with a single, gradual topic drift.

None of these is derived from a generative model of topic structure. They are heuristics validated empirically on retrieval quality.

Relation to classical text segmentation

Topic segmentation predates embeddings. Hearst (1997), TextTiling: Segmenting Text into Multi-Paragraph Subtopic Passages, Computational Linguistics 23(1), used lexical cohesion for the same kind of boundary. It compared shared vocabulary between adjacent blocks of text, via cosine similarity over word-count vectors, instead of dense embeddings. The method here is TextTiling with e_i swapped from a sparse bag-of-words vector to a dense contextual one. Cutting where local cohesion drops is a decades-old idea.

Cost

For a document of n sentences and embedding dimension d, computing all embeddings costs one forward pass per sentence, batchable to O(n) calls amortised over batch size. Pairwise similarity across consecutive sentences only, not all pairs, costs O(n * d), avoiding the O(n^2 * d) cost of comparing every sentence to every other sentence. This locality is what keeps semantic chunking affordable even on long documents.

Honest limitations

This lesson's threshold rules are practitioner conventions, popularised in retrieval-pipeline libraries and blog posts rather than established through peer-reviewed comparison. No controlled study is widely cited showing one threshold rule reliably beats the others across document types. Treat the choice as a tunable hyperparameter, evaluated on your own retrieval metrics — see Building a retrieval eval set — not as a settled result.

Key references

  • Hearst, M. (1997). TextTiling: Segmenting Text into Multi-Paragraph Subtopic Passages. Computational Linguistics 23(1). The classical, embedding-free ancestor of this technique.
  • Karpukhin, V. et al. (2020). Dense Passage Retrieval for Open-Domain Question Answering. arXiv:2004.04906 — establishes the dense-embedding similarity this lesson's breakpoint detection depends on.
  • Gao, Y. et al. (2023). Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997

Current state and open problems

Semantic chunking is a pre-processing heuristic, not a learned system, and it inherits every limitation of the sentence embedding model it uses. A model with weak domain coverage — heavy jargon, code, or a low-resource language — produces unreliable similarity scores, and the breakpoints follow.

Learned segmentation, where a model is trained end-to-end to predict good chunk boundaries for a specific retrieval task, remains a research direction rather than a deployed default. The gap between "topically coherent" and "useful for this specific downstream question" is still bridged by evaluation against real queries, not by the segmentation method alone.

What to learn next