ColBERT and late interaction
Instead of squashing a whole document into one vector, keep one vector per word and match each query word to its single best-matching document word — a genuine middle ground between bi-encoders and cross-encoders.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Late interaction keeps one vector per word, instead of squashing a whole document into a single vector. It matches word to word, then adds up a score.
Picture checking a student's answer against a model answer sheet, two different ways. One teacher reads the whole essay, forms a general impression, and gives a single overall grade. Another goes phrase by phrase. She checks whether this specific claim matches any claim in the model answer. She does that for every phrase, then adds up the results. The second teacher's grade is slower to produce. It catches specific matches and mismatches the first teacher's "vibe" grade can blur together.
The bi-encoder from earlier in this section works like the first teacher. It squashes a whole document into one vector, an overall impression. ColBERT's late interaction approach works like the second teacher instead. It keeps one vector per word, and compares query words against document words individually. The results combine into one final score.
Why it exists
A single pooled vector is efficient but lossy. Cramming everything a sentence means into one fixed-length vector inevitably blurs some detail. A cross-encoder avoids that loss by reading the query and document together. It pays for that with a full model pass per document, on every single search. Nothing there is precomputable in advance.
ColBERT, introduced by Khattab & Zaharia in 2020, is built to sit between the two. It keeps a whole set of per-word vectors for every document. Like a bi-encoder's vector, they are precomputable and storable ahead of time. It then compares them word-by-word at search time, catching detail a single pooled vector would have lost.
How it works
query: "waterproof speaker" document: "bluetooth speaker, waterproof
and portable, for outdoor parties"
BI-ENCODER: squash query into ONE vector, squash document into ONE vector,
compare those two vectors -- fine detail about individual words gets blended away
LATE INTERACTION (ColBERT-style): keep a vector for EVERY word
query words: [water] [proof] [speaker]
doc words: [bluetooth] [speaker] [waterproof] [and] [portable] ...
for "water": best match in the doc is "waterproof" -> high score
for "speaker": best match in the doc is "speaker" -> high score
|
v
add up each query word's BEST match = final scoreThis word-by-word best-match-and-sum operation is called MaxSim. Every query word finds its own single best partner in the document. The final score is the total of those best partners. A single pooled vector would have blended all of that detail away.
A real example you have seen
Sometimes a document deserves a high rank because of one exactly-matching phrase. The rest of it may be only loosely related. Late-interaction search is built to behave exactly that way. A single pooled-vector comparison tends to average that one strong match away, against the rest of the document.
Remember this
- Late interaction keeps one vector per word, not one vector per document, unlike a standard bi-encoder.
- MaxSim matches each query word to its single best-matching document word, then sums those best matches into a final score.
- It sits between a bi-encoder (fast, coarser) and a cross-encoder (slow, sharpest) in both cost and accuracy.
What to learn next
- Learned sparse retrieval with SPLADE — a different way to get past a single pooled vector's blurriness, using sparse, word-aligned scores instead.
- Bi-encoders vs cross-encoders — the two endpoints late interaction sits between.
- Reranking the top 50 — where a late-interaction model can slot into an existing pipeline.
Developer — Code and libraries.
This demonstrates the MaxSim mechanism itself, using an off-the-shelf sentence-embedding model's per-token output. This is not the real, specially-trained ColBERT checkpoint — it illustrates how the scoring works, using a general-purpose model that was never trained for late interaction specifically. Treat the numbers below as a demonstration of the mechanism, not as retrieval-quality scores.
Setup
pip install sentence-transformers torchOutputs verified with sentence-transformers 5.4.1 and torch 2.5.1 on CPU.
Comparing a single pooled score against a MaxSim score
import torch
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
tokenizer = model.tokenizer
query = "waterproof speaker"
doc_a = "bluetooth speaker, waterproof and portable, great for outdoor parties"
doc_b = "wireless earbuds with noise cancellation and 24 hour battery life"
def token_embeddings(text):
out = model.encode(text, output_value="token_embeddings", convert_to_tensor=True)
out = torch.nn.functional.normalize(out, dim=-1)
toks = tokenizer.convert_ids_to_tokens(tokenizer(text)["input_ids"])
return toks, out
def maxsim_score(query_text, doc_text):
q_toks, q_emb = token_embeddings(query_text)
d_toks, d_emb = token_embeddings(doc_text)
sim = q_emb @ d_emb.T # every query token against every doc token
best_per_query_token = sim.max(dim=1).values
return best_per_query_token.sum().item(), q_toks, best_per_query_token
def single_vector_score(query_text, doc_text):
q = model.encode(query_text, normalize_embeddings=True)
d = model.encode(doc_text, normalize_embeddings=True)
return float(q @ d)
for label, doc in [("doc A (waterproof speaker)", doc_a), ("doc B (unrelated earbuds)", doc_b)]:
late_score, q_toks, per_tok = maxsim_score(query, doc)
single_score = single_vector_score(query, doc)
print(f"{label}")
print(f" single-vector cosine similarity: {single_score:.3f}")
print(f" ColBERT-style MaxSim score: {late_score:.3f}")
for tok, s in zip(q_toks, per_tok.tolist()):
print(f" query token {tok!r:>12} -> best match in doc: {s:.3f}")
print()doc A (waterproof speaker)
single-vector cosine similarity: 0.687
ColBERT-style MaxSim score: 4.061
query token '[CLS]' -> best match in doc: 0.814
query token 'water' -> best match in doc: 0.901
query token '##proof' -> best match in doc: 0.843
query token 'speaker' -> best match in doc: 0.906
query token '[SEP]' -> best match in doc: 0.597
doc B (unrelated earbuds)
single-vector cosine similarity: 0.280
ColBERT-style MaxSim score: 1.650
query token '[CLS]' -> best match in doc: 0.655
query token 'water' -> best match in doc: 0.221
query token '##proof' -> best match in doc: 0.238
query token 'speaker' -> best match in doc: 0.257
query token '[SEP]' -> best match in doc: 0.279Both scoring methods agree doc A is the better match. The token-level breakdown shows exactly why: "water" and "speaker" each find an almost perfect partner (0.90+) somewhere in doc A, and a weak one (around 0.22–0.26) in doc B — visibility a single pooled cosine similarity number cannot offer.
The walkthrough
output_value="token_embeddings" returns one vector per token, before pooling. A standard model.encode() call runs this same per-token output through a pooling step (usually mean pooling) to collapse it into one vector. Skipping that pooling step is what turns an ordinary bi-encoder into a late-interaction scorer.
sim.max(dim=1) is the entire MaxSim operation. For each query token's row in the similarity matrix, take the single highest value across every document token's column — the query token's best possible partner anywhere in the document, regardless of position.
[CLS] and [SEP] tokens are noise here, left in for transparency. Real ColBERT implementations typically mask out special tokens (and often punctuation) before scoring — they are shown unmasked above so nothing about the raw mechanism is hidden.
Why this is an illustration, not real ColBERT. The real ColBERT model is trained end-to-end with a ranking loss specifically shaping its token embeddings to work well under MaxSim scoring — including a lightweight linear projection down to a smaller per-token dimension for storage efficiency. all-MiniLM-L6-v2, used here, was trained for pooled-sentence similarity, not late interaction; its token-level vectors happen to demonstrate the mechanism well, but should not be read as evidence of real ColBERT-level retrieval quality.
Common mistakes
Forgetting to mask special and punctuation tokens in a real implementation. Left unmasked, as shown above, they contribute noise to every score roughly equally — usually harmless for ranking (since it affects every candidate similarly), but wasted signal a real implementation should strip out.
Assuming late interaction is "free" compared to a cross-encoder. It is far cheaper — document token vectors are precomputed, the same way a bi-encoder's pooled vector is — but storing many vectors per document instead of one is a real memory and index-complexity cost, addressed directly in ColBERTv2's compression work, covered in the researcher block.
Comparing MaxSim scores across queries of different lengths without normalising. The MaxSim sum in this implementation grows with the number of query tokens — a five-word query and a two-word query are not on a directly comparable scale unless you normalise by query length.
Try it yourself
Change the query to "loud portable music" — no exact word overlap with either document at all — and re-run. Compare how much the single-vector cosine score moves versus how the per-token breakdown explains which specific words are driving the (weaker, since there's no exact overlap this time) match to doc A.
What to learn next
- Learned sparse retrieval with SPLADE — a different way to get past a single pooled vector's blurriness, using sparse, word-aligned scores instead.
- Bi-encoders vs cross-encoders — the two endpoints late interaction sits between.
- Reranking the top 50 — where a late-interaction model can slot into an existing pipeline.
Researcher — Mathematics and papers.
The scoring function
Following Khattab & Zaharia (2020), given query token embeddings E_q = {q_1, ..., q_m} and document token embeddings E_d = {d_1, ..., d_n}, the ColBERT relevance score is:
Score(q, d) = sum over i=1..m of max over j=1..n of (q_i . d_j)Where each query token independently finds its best-matching document token via a dot product (equivalent to cosine similarity for normalised vectors), and the final score sums these per-token maxima. This is the MaxSim operator demonstrated directly in the developer block.
Both E_q and E_d are produced by a shared BERT encoder, but — critically — E_d can be computed and stored before any query arrives, exactly like a bi-encoder's document vectors, while the fine-grained token-level matching happens only at query time, "late" in the pipeline. This is the origin of the name late interaction: interaction between query and document tokens is deferred to scoring time, unlike a cross-encoder's "early" interaction (full attention from layer one) or a standard bi-encoder's complete absence of interaction (a single pooled score, no token-level comparison at all).
Complexity
For a candidate set of N documents with average length n tokens and a query of m tokens, MaxSim scoring costs O(N * m * n * d) for embedding dimension d — linear in the number of candidates, unlike a cross-encoder's cost, which requires a full transformer forward pass per candidate. This makes late interaction dramatically cheaper than cross-encoder reranking at the candidate-set sizes reranking typically operates on, while retaining meaningfully more token-level precision than a single pooled bi-encoder vector.
ColBERTv2 and the storage problem
The direct cost of late interaction is storage: keeping one vector per token, rather than one per document, multiplies index size by roughly the average document length in tokens. ColBERTv2 (Santhanam, Khattab, Saad-Falcon, Potts & Zaharia, 2021) addresses this with residual compression — each token vector is expressed as a small offset from its nearest centroid in a learned codebook, quantized aggressively, shrinking the index by a large factor with minimal quality loss, and making late interaction practical at genuinely large corpus scale.
Key references
- Khattab, O. & Zaharia, M. (2020). ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. arXiv:2004.12832
- Santhanam, K., Khattab, O., Saad-Falcon, J., Potts, C. & Zaharia, M. (2021). ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction. arXiv:2112.01488
Current state
Late interaction remains an active middle ground in the retrieval literature between the speed of bi-encoders and the accuracy of cross-encoders, and ColBERTv2's compression work has made it practical to deploy at real production scale rather than only as a research curiosity. It is used both as a standalone first-stage retriever (searching MaxSim scores directly against a compressed token index) and as a reranking stage, occupying a genuinely distinct cost-accuracy point from either the bi-encoder or cross-encoder covered earlier in this section.
What to learn next
- Learned sparse retrieval with SPLADE — a different way to get past a single pooled vector's blurriness, using sparse, word-aligned scores instead.
- Bi-encoders vs cross-encoders — the two endpoints late interaction sits between.
- Reranking the top 50 — where a late-interaction model can slot into an existing pipeline.