Semantic Search and Reranking

ColBERT and late interaction

Instead of squashing a whole document into one vector, keep one vector per word and match each query word to its single best-matching document word — a genuine middle ground between bi-encoders and cross-encoders.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Late interaction keeps one vector per word, instead of squashing a whole document into a single vector. It matches word to word, then adds up a score.

Picture checking a student's answer against a model answer sheet, two different ways. One teacher reads the whole essay, forms a general impression, and gives a single overall grade. Another goes phrase by phrase. She checks whether this specific claim matches any claim in the model answer. She does that for every phrase, then adds up the results. The second teacher's grade is slower to produce. It catches specific matches and mismatches the first teacher's "vibe" grade can blur together.

The bi-encoder from earlier in this section works like the first teacher. It squashes a whole document into one vector, an overall impression. ColBERT's late interaction approach works like the second teacher instead. It keeps one vector per word, and compares query words against document words individually. The results combine into one final score.

Why it exists

A single pooled vector is efficient but lossy. Cramming everything a sentence means into one fixed-length vector inevitably blurs some detail. A cross-encoder avoids that loss by reading the query and document together. It pays for that with a full model pass per document, on every single search. Nothing there is precomputable in advance.

ColBERT, introduced by Khattab & Zaharia in 2020, is built to sit between the two. It keeps a whole set of per-word vectors for every document. Like a bi-encoder's vector, they are precomputable and storable ahead of time. It then compares them word-by-word at search time, catching detail a single pooled vector would have lost.

How it works

query: "waterproof speaker"                document: "bluetooth speaker, waterproof
                                             and portable, for outdoor parties"

BI-ENCODER: squash query into ONE vector, squash document into ONE vector,
compare those two vectors -- fine detail about individual words gets blended away

LATE INTERACTION (ColBERT-style): keep a vector for EVERY word
  query words:    [water] [proof] [speaker]
  doc words:      [bluetooth] [speaker] [waterproof] [and] [portable] ...

  for "water":    best match in the doc is "waterproof"  -> high score
  for "speaker":  best match in the doc is "speaker"     -> high score
                          |
                          v
              add up each query word's BEST match  =  final score

This word-by-word best-match-and-sum operation is called MaxSim. Every query word finds its own single best partner in the document. The final score is the total of those best partners. A single pooled vector would have blended all of that detail away.

A real example you have seen

Sometimes a document deserves a high rank because of one exactly-matching phrase. The rest of it may be only loosely related. Late-interaction search is built to behave exactly that way. A single pooled-vector comparison tends to average that one strong match away, against the rest of the document.

Remember this

  • Late interaction keeps one vector per word, not one vector per document, unlike a standard bi-encoder.
  • MaxSim matches each query word to its single best-matching document word, then sums those best matches into a final score.
  • It sits between a bi-encoder (fast, coarser) and a cross-encoder (slow, sharpest) in both cost and accuracy.

What to learn next

Developer — Code and libraries.

This demonstrates the MaxSim mechanism itself, using an off-the-shelf sentence-embedding model's per-token output. This is not the real, specially-trained ColBERT checkpoint — it illustrates how the scoring works, using a general-purpose model that was never trained for late interaction specifically. Treat the numbers below as a demonstration of the mechanism, not as retrieval-quality scores.

Setup

bash
pip install sentence-transformers torch

Outputs verified with sentence-transformers 5.4.1 and torch 2.5.1 on CPU.

Comparing a single pooled score against a MaxSim score

maxsim_demo.py
import torch
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("all-MiniLM-L6-v2")
tokenizer = model.tokenizer

query = "waterproof speaker"
doc_a = "bluetooth speaker, waterproof and portable, great for outdoor parties"
doc_b = "wireless earbuds with noise cancellation and 24 hour battery life"


def token_embeddings(text):
    out = model.encode(text, output_value="token_embeddings", convert_to_tensor=True)
    out = torch.nn.functional.normalize(out, dim=-1)
    toks = tokenizer.convert_ids_to_tokens(tokenizer(text)["input_ids"])
    return toks, out


def maxsim_score(query_text, doc_text):
    q_toks, q_emb = token_embeddings(query_text)
    d_toks, d_emb = token_embeddings(doc_text)
    sim = q_emb @ d_emb.T                      # every query token against every doc token
    best_per_query_token = sim.max(dim=1).values
    return best_per_query_token.sum().item(), q_toks, best_per_query_token


def single_vector_score(query_text, doc_text):
    q = model.encode(query_text, normalize_embeddings=True)
    d = model.encode(doc_text, normalize_embeddings=True)
    return float(q @ d)


for label, doc in [("doc A (waterproof speaker)", doc_a), ("doc B (unrelated earbuds)", doc_b)]:
    late_score, q_toks, per_tok = maxsim_score(query, doc)
    single_score = single_vector_score(query, doc)
    print(f"{label}")
    print(f"  single-vector cosine similarity: {single_score:.3f}")
    print(f"  ColBERT-style MaxSim score:       {late_score:.3f}")
    for tok, s in zip(q_toks, per_tok.tolist()):
        print(f"      query token {tok!r:>12} -> best match in doc: {s:.3f}")
    print()
Output
doc A (waterproof speaker)
  single-vector cosine similarity: 0.687
  ColBERT-style MaxSim score:       4.061
      query token      '[CLS]' -> best match in doc: 0.814
      query token      'water' -> best match in doc: 0.901
      query token    '##proof' -> best match in doc: 0.843
      query token    'speaker' -> best match in doc: 0.906
      query token      '[SEP]' -> best match in doc: 0.597

doc B (unrelated earbuds)
  single-vector cosine similarity: 0.280
  ColBERT-style MaxSim score:       1.650
      query token      '[CLS]' -> best match in doc: 0.655
      query token      'water' -> best match in doc: 0.221
      query token    '##proof' -> best match in doc: 0.238
      query token    'speaker' -> best match in doc: 0.257
      query token      '[SEP]' -> best match in doc: 0.279

Both scoring methods agree doc A is the better match. The token-level breakdown shows exactly why: "water" and "speaker" each find an almost perfect partner (0.90+) somewhere in doc A, and a weak one (around 0.22–0.26) in doc B — visibility a single pooled cosine similarity number cannot offer.

The walkthrough

output_value="token_embeddings" returns one vector per token, before pooling. A standard model.encode() call runs this same per-token output through a pooling step (usually mean pooling) to collapse it into one vector. Skipping that pooling step is what turns an ordinary bi-encoder into a late-interaction scorer.

sim.max(dim=1) is the entire MaxSim operation. For each query token's row in the similarity matrix, take the single highest value across every document token's column — the query token's best possible partner anywhere in the document, regardless of position.

[CLS] and [SEP] tokens are noise here, left in for transparency. Real ColBERT implementations typically mask out special tokens (and often punctuation) before scoring — they are shown unmasked above so nothing about the raw mechanism is hidden.

Why this is an illustration, not real ColBERT. The real ColBERT model is trained end-to-end with a ranking loss specifically shaping its token embeddings to work well under MaxSim scoring — including a lightweight linear projection down to a smaller per-token dimension for storage efficiency. all-MiniLM-L6-v2, used here, was trained for pooled-sentence similarity, not late interaction; its token-level vectors happen to demonstrate the mechanism well, but should not be read as evidence of real ColBERT-level retrieval quality.

Common mistakes

Forgetting to mask special and punctuation tokens in a real implementation. Left unmasked, as shown above, they contribute noise to every score roughly equally — usually harmless for ranking (since it affects every candidate similarly), but wasted signal a real implementation should strip out.

Assuming late interaction is "free" compared to a cross-encoder. It is far cheaper — document token vectors are precomputed, the same way a bi-encoder's pooled vector is — but storing many vectors per document instead of one is a real memory and index-complexity cost, addressed directly in ColBERTv2's compression work, covered in the researcher block.

Comparing MaxSim scores across queries of different lengths without normalising. The MaxSim sum in this implementation grows with the number of query tokens — a five-word query and a two-word query are not on a directly comparable scale unless you normalise by query length.

Try it yourself

Change the query to "loud portable music" — no exact word overlap with either document at all — and re-run. Compare how much the single-vector cosine score moves versus how the per-token breakdown explains which specific words are driving the (weaker, since there's no exact overlap this time) match to doc A.

What to learn next

Researcher — Mathematics and papers.

The scoring function

Following Khattab & Zaharia (2020), given query token embeddings E_q = {q_1, ..., q_m} and document token embeddings E_d = {d_1, ..., d_n}, the ColBERT relevance score is:

text
Score(q, d) = sum over i=1..m of  max over j=1..n of (q_i . d_j)

Where each query token independently finds its best-matching document token via a dot product (equivalent to cosine similarity for normalised vectors), and the final score sums these per-token maxima. This is the MaxSim operator demonstrated directly in the developer block.

Both E_q and E_d are produced by a shared BERT encoder, but — critically — E_d can be computed and stored before any query arrives, exactly like a bi-encoder's document vectors, while the fine-grained token-level matching happens only at query time, "late" in the pipeline. This is the origin of the name late interaction: interaction between query and document tokens is deferred to scoring time, unlike a cross-encoder's "early" interaction (full attention from layer one) or a standard bi-encoder's complete absence of interaction (a single pooled score, no token-level comparison at all).

Complexity

For a candidate set of N documents with average length n tokens and a query of m tokens, MaxSim scoring costs O(N * m * n * d) for embedding dimension d — linear in the number of candidates, unlike a cross-encoder's cost, which requires a full transformer forward pass per candidate. This makes late interaction dramatically cheaper than cross-encoder reranking at the candidate-set sizes reranking typically operates on, while retaining meaningfully more token-level precision than a single pooled bi-encoder vector.

ColBERTv2 and the storage problem

The direct cost of late interaction is storage: keeping one vector per token, rather than one per document, multiplies index size by roughly the average document length in tokens. ColBERTv2 (Santhanam, Khattab, Saad-Falcon, Potts & Zaharia, 2021) addresses this with residual compression — each token vector is expressed as a small offset from its nearest centroid in a learned codebook, quantized aggressively, shrinking the index by a large factor with minimal quality loss, and making late interaction practical at genuinely large corpus scale.

Key references

  • Khattab, O. & Zaharia, M. (2020). ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. arXiv:2004.12832
  • Santhanam, K., Khattab, O., Saad-Falcon, J., Potts, C. & Zaharia, M. (2021). ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction. arXiv:2112.01488

Current state

Late interaction remains an active middle ground in the retrieval literature between the speed of bi-encoders and the accuracy of cross-encoders, and ColBERTv2's compression work has made it practical to deploy at real production scale rather than only as a research curiosity. It is used both as a standalone first-stage retriever (searching MaxSim scores directly against a compressed token index) and as a reranking stage, occupying a genuinely distinct cost-accuracy point from either the bi-encoder or cross-encoder covered earlier in this section.

What to learn next