Semantic chunking
Semantic chunking splits a document where its topic actually shifts, using meaning rather than a fixed character count to decide where one chunk ends and the next begins.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Semantic chunking splits a document where the topic changes, not at a fixed character count.
Think about sorting a mixed box of family photographs into piles. Birthday photos go in one pile, a wedding in another, a holiday trip in a third. You are not counting photos into equal stacks. You are grouping them by what each one is actually about.
Why it exists
The chunking methods from the last lesson cut by character count or by sentence boundary. Both ignore meaning entirely.
That causes a real problem. Two unrelated sentences can land in the same chunk, only because they both fit under the size limit. A search system then retrieves that chunk for a question about one topic. It drags in irrelevant text from the other.
Semantic chunking fixes this by measuring how similar each sentence is to the one before it. A large drop in similarity signals a topic change — a good place to cut.
How it works
Each sentence gets turned into a list of numbers, called an embedding, that captures its meaning. Sentences about similar things get similar numbers.
Sentence: [about interest rates] [about interest rates] [about monsoon rainfall]
Similarity
to next: 0.83 0.15 <- big drop, cut here
0.45
Result:
Chunk A: the two interest-rate sentences, kept together
Chunk B: the rainfall sentence, starts a new chunkA small drop in similarity between two sentences means they likely belong together. A big drop means the topic moved on, and that is where the chunk boundary goes.
Where you have already seen it
- Podcast and video apps with auto-chapters. A single episode gets split into named sections, often based on where the topic in the transcript shifts.
- Meeting-notes tools. Notes from an hour-long call get grouped under separate headings per discussion topic, not by the clock.
- Long WhatsApp or email threads, summarised. A summary tool that separates "budget discussion" from "hiring update" within one thread is doing something similar.
Remember this
- Semantic chunking measures meaning, using embeddings, rather than counting characters alone.
- A big drop in sentence-to-sentence similarity marks a good place to cut.
- It costs more compute than fixed-size chunking, since every sentence needs an embedding first.
What to learn next
- Chunking strategies compared — the simpler methods this one improves on.
- Embeddings — what those "numbers that capture meaning" actually are.
- Giving each chunk its context back — a companion fix for chunks that lose their surrounding topic.
Developer — Code and libraries.
This builds real sentence embeddings with a small pretrained model, then finds the breakpoint where topic similarity drops.
Setup
pip install transformers torchThe first run downloads sentence-transformers/all-MiniLM-L6-v2, about 90 MB. It caches locally after that.
Finding the semantic breakpoint
import re
import torch
from transformers import AutoTokenizer, AutoModel
doc = (
"The Reserve Bank of India sets the repo rate. "
"The repo rate is the interest rate at which the RBI lends to commercial banks. "
"A higher repo rate makes loans more expensive for everyone. "
"Monsoon rainfall this year was above average across most states. "
"Good rainfall usually means a strong harvest for rice and wheat farmers. "
"A strong harvest tends to keep food prices low in local markets."
)
sentences = re.split(r'(?<=[.!?])\s+', doc.strip())
tok = AutoTokenizer.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model = AutoModel.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model.eval()
enc = tok(sentences, padding=True, truncation=True, return_tensors="pt")
with torch.no_grad():
out = model(**enc)
# mean-pool the token embeddings into one vector per sentence
mask = enc["attention_mask"].unsqueeze(-1).float()
pooled = (out.last_hidden_state * mask).sum(1) / mask.sum(1)
pooled = torch.nn.functional.normalize(pooled, dim=1)
sims = (pooled[:-1] * pooled[1:]).sum(dim=1) # cosine similarity, consecutive pairs
for i, s in enumerate(sentences):
tail = f" similarity to next: {sims[i]:.3f}" if i < len(sims) else ""
print(f"{i}: {s}{tail}")
threshold = sims.mean() - sims.std()
breakpoints = [i for i, s in enumerate(sims) if s < threshold]
print(f"\nbreakpoint threshold: {threshold:.3f}")
print(f"split after sentence index: {breakpoints}")0: The Reserve Bank of India sets the repo rate. similarity to next: 0.829 1: The repo rate is the interest rate at which the RBI lends to commercial banks. similarity to next: 0.635 2: A higher repo rate makes loans more expensive for everyone. similarity to next: 0.149 3: Monsoon rainfall this year was above average across most states. similarity to next: 0.398 4: Good rainfall usually means a strong harvest for rice and wheat farmers. similarity to next: 0.448 5: A strong harvest tends to keep food prices low in local markets. breakpoint threshold: 0.236 split after sentence index: [2]
Model weights are pinned, so the numbers above are exact for this model version — a different transformers release could shift them slightly at the third decimal.
Line by line
pooled = (out.last_hidden_state * mask).sum(1) / mask.sum(1) builds one vector per sentence by averaging its token vectors, ignoring padding. This is mean pooling, the same operation the sentence-transformers library performs internally for this model family.
normalize(pooled, dim=1) scales each vector to length 1. After that, a dot product between two vectors equals their cosine similarity, which is why the next line can skip a separate cosine-similarity function.
sims = (pooled[:-1] * pooled[1:]).sum(dim=1) multiplies every sentence's vector against the next sentence's vector, element-wise, then sums. That is a dot product computed for every consecutive pair at once, with no loop.
The threshold, mean - std, is one common rule of thumb: cut wherever similarity drops more than one standard deviation below the document's average. It found the real topic boundary here — between "makes loans more expensive" and "monsoon rainfall" — with no keyword rule telling it where to look.
Common mistakes
Running this on a two- or three-sentence document. Standard deviation over three numbers is close to meaningless, and the threshold becomes unstable. Semantic chunking needs a reasonable number of sentences to have a stable notion of "average similarity."
Assuming the threshold formula is universal. mean - std is one convention among several. Some implementations use a percentile cutoff instead, such as the 5th percentile of all pairwise similarities. Test both on your own documents.
Re-embedding every sentence for every new document, live, in a user-facing request. Embedding is cheap per sentence but adds up. Production systems chunk documents once, offline, and store the result — not on every search.
Forgetting this still needs a maximum chunk size. A long run of highly similar sentences never triggers a break on its own. Pair semantic chunking with a hard size cap, so one topic that runs for ten pages does not become one giant chunk.
Try it yourself
Add a fourth topic to doc — one or two more sentences about something else entirely, like a cricket score. Re-run the script.
You should see a new similarity drop appear right before your added sentences, and a new entry in breakpoints. The threshold is recomputed from the whole document each time, so adding text changes the mean and standard deviation slightly, not only where you added the sentence.
What to learn next
- Chunking strategies compared — the fixed-rule baseline this method replaces.
- Sentence transformers — more on the model family used here.
- Keyword vs semantic search — the same embedding idea, applied to search instead of splitting.
Researcher — Mathematics and papers.
Formal definition
Given a sequence of sentence embeddings e_1, ..., e_n, define the adjacent similarity series:
sim_i = cos(e_i, e_{i+1}) = (e_i . e_{i+1}) / (||e_i|| * ||e_{i+1}||), i = 1..n-1e_iis the pooled embedding of sentencei.cos(.,.)is cosine similarity, in[-1, 1]for arbitrary vectors and typically[0, 1]for sentence embeddings from models trained with a contrastive objective.
A breakpoint is any index i where sim_i falls below a threshold tau. The threshold is a design choice, not derived from theory. Common conventions:
- Standard-deviation rule:
tau = mean(sim) - k * std(sim), withkaround 1. - Percentile rule:
tauset to thep-th percentile of thesimdistribution,ptypically between 5 and 25. - Local-minimum rule: cut only at sentences that are a local minimum in
sim, which avoids over-splitting a document with a single, gradual topic drift.
None of these is derived from a generative model of topic structure. They are heuristics validated empirically on retrieval quality.
Relation to classical text segmentation
Topic segmentation predates embeddings. Hearst (1997), TextTiling: Segmenting Text into Multi-Paragraph Subtopic Passages, Computational Linguistics 23(1), used lexical cohesion for the same kind of boundary. It compared shared vocabulary between adjacent blocks of text, via cosine similarity over word-count vectors, instead of dense embeddings. The method here is TextTiling with e_i swapped from a sparse bag-of-words vector to a dense contextual one. Cutting where local cohesion drops is a decades-old idea.
Cost
For a document of n sentences and embedding dimension d, computing all embeddings costs one forward pass per sentence, batchable to O(n) calls amortised over batch size. Pairwise similarity across consecutive sentences only, not all pairs, costs O(n * d), avoiding the O(n^2 * d) cost of comparing every sentence to every other sentence. This locality is what keeps semantic chunking affordable even on long documents.
Honest limitations
This lesson's threshold rules are practitioner conventions, popularised in retrieval-pipeline libraries and blog posts rather than established through peer-reviewed comparison. No controlled study is widely cited showing one threshold rule reliably beats the others across document types. Treat the choice as a tunable hyperparameter, evaluated on your own retrieval metrics — see Building a retrieval eval set — not as a settled result.
Key references
- Hearst, M. (1997). TextTiling: Segmenting Text into Multi-Paragraph Subtopic Passages. Computational Linguistics 23(1). The classical, embedding-free ancestor of this technique.
- Karpukhin, V. et al. (2020). Dense Passage Retrieval for Open-Domain Question Answering. arXiv:2004.04906 — establishes the dense-embedding similarity this lesson's breakpoint detection depends on.
- Gao, Y. et al. (2023). Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997
Current state and open problems
Semantic chunking is a pre-processing heuristic, not a learned system, and it inherits every limitation of the sentence embedding model it uses. A model with weak domain coverage — heavy jargon, code, or a low-resource language — produces unreliable similarity scores, and the breakpoints follow.
Learned segmentation, where a model is trained end-to-end to predict good chunk boundaries for a specific retrieval task, remains a research direction rather than a deployed default. The gap between "topically coherent" and "useful for this specific downstream question" is still bridged by evaluation against real queries, not by the segmentation method alone.
What to learn next
- Chunking strategies compared — the rule-based baseline.
- Contextual vs static embeddings — how the embeddings behind this method are built.
- Bi-encoders vs cross-encoders — a deeper look at the model family used to embed sentences.