Bi-encoders vs cross-encoders
A bi-encoder scores a query against a document by comparing two vectors computed separately. A cross-encoder reads both together at once — slower, but noticeably sharper.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A bi-encoder judges a match by comparing two things separately. A cross-encoder judges a match by looking at both things together, at the same time.
Think about hiring for a job. A recruiter skimming resumes compares each one, on its own, against a checklist of what the role needs. It is fast, and it happens without the candidate and the job description ever being in the same room. A real interview is different. The interviewer and the candidate sit together, and questions and answers shape each other in real time. The interview takes far longer, but it catches nuance the resume skim misses.
Bi-encoders work like resume screening. The query and every document each get turned into an embedding independently. A match is only a distance between two already-finished vectors. Cross-encoders work like the interview instead. The query and one document are fed into the model together. Every word of one can influence how every word of the other is understood. One relevance score comes out at the end.
Why it exists
Semantic search needs to compare a query against potentially millions of documents. Doing that with a cross-encoder means running the model millions of times, once per document, for every single search. That is far too slow to use directly at that scale.
Bi-encoders solve the scale problem. Every document's embedding can be computed once, ahead of time, and stored. At search time, only the query needs encoding. Comparing it against a million stored vectors is a fast, cheap operation. The trade-off is real, though. A bi-encoder never lets the query and document "look at" each other while forming their vectors. It can miss subtler matches a cross-encoder would catch.
How it works
BI-ENCODER CROSS-ENCODER
query -> [model] -> vector Q query + document, TOGETHER
doc 1 -> [model] -> vector D1 |
doc 2 -> [model] -> vector D2 v
... (computed once, stored) [model reads both at once]
|
compare Q against every D with v
one fast similarity calculation a single relevance score
FAST at search time. SLOW -- one full model pass
Documents encoded once, reused forever. PER document, every single search.The two are not really competitors. They sit on a speed-versus-accuracy trade-off. Production search systems typically use both — a pattern covered fully in the next lesson.
A real example you have seen
Search a huge product catalogue and results appear instantly. A bi-encoder is doing the heavy lifting there. It compares your query against millions of pre-computed vectors in a fraction of a second. A shopping site may then show a small, unusually well-ordered set of "best matches" at the very top. A cross-encoder is often quietly re-checking only that short list.
Remember this
- A bi-encoder encodes the query and each document separately — fast, and reusable across every future search.
- A cross-encoder reads the query and one document together — slower, but catches nuance a bi-encoder alone misses.
- Bi-encoders scale to millions of documents; cross-encoders are typically reserved for a short list, not a whole catalogue.
What to learn next
- Reranking the top 50 — the full two-stage pipeline this lesson previewed.
- ColBERT and late interaction — a third architecture between bi-encoders and cross-encoders.
- Attention — the mechanism giving a cross-encoder its query-document interaction.
Developer — Code and libraries.
Setup
pip install sentence-transformersTwo models download on first run: all-MiniLM-L6-v2 (the bi-encoder, ~90 MB) and cross-encoder/ms-marco-MiniLM-L-6-v2 (the cross-encoder, ~90 MB). Outputs verified with sentence-transformers 5.4.1 on CPU.
Same query, same five candidates, two different judges
import numpy as np
from sentence_transformers import SentenceTransformer, CrossEncoder
products = [
"Wireless earbuds with noise cancellation and 24 hour battery life",
"Over-ear headphones with deep bass and a foldable design",
"Bluetooth speaker, waterproof and portable, great for outdoor parties",
"Laptop with 16GB RAM and fast SSD storage, built for programming",
"Budget laptop for students, lightweight with long battery life",
"Mechanical keyboard with RGB lighting and tactile switches",
"Wireless mouse with an ergonomic shape that works on any surface",
"Smartphone with a triple camera and all-day battery",
"Budget smartphone with a big screen for watching videos",
"Smartwatch that tracks your heart rate and sleep",
"Fitness band with a step counter and water resistance",
"External hard drive with 1TB of storage and fast USB-C transfer",
"Portable power bank that charges your phone twice over",
"Over-ear headphones built for long flights, with active noise cancelling",
"4K webcam for video calls and live streaming",
]
query = "something to bring peace and quiet on a long flight"
bi_encoder = SentenceTransformer("all-MiniLM-L6-v2")
doc_vecs = bi_encoder.encode(products, normalize_embeddings=True)
q_vec = bi_encoder.encode([query], normalize_embeddings=True)[0]
bi_scores = doc_vecs @ q_vec
bi_top5 = np.argsort(bi_scores)[::-1][:5]
print("bi-encoder top 5 (cosine similarity, precomputed doc vectors):")
for i in bi_top5:
print(f" {bi_scores[i]:.3f} {products[i]}")
cross_encoder = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")
pairs = [(query, products[i]) for i in bi_top5]
cross_scores = cross_encoder.predict(pairs)
order = np.argsort(cross_scores)[::-1]
print("\nsame 5 candidates, re-scored by the cross-encoder (query+doc read together):")
for j in order:
i = bi_top5[j]
print(f" {cross_scores[j]:.3f} {products[i]}")bi-encoder top 5 (cosine similarity, precomputed doc vectors): 0.425 Over-ear headphones built for long flights, with active noise cancelling 0.237 Bluetooth speaker, waterproof and portable, great for outdoor parties 0.188 Budget laptop for students, lightweight with long battery life 0.165 Wireless earbuds with noise cancellation and 24 hour battery life 0.153 Over-ear headphones with deep bass and a foldable design same 5 candidates, re-scored by the cross-encoder (query+doc read together): -3.966 Over-ear headphones built for long flights, with active noise cancelling -10.847 Bluetooth speaker, waterproof and portable, great for outdoor parties -10.969 Wireless earbuds with noise cancellation and 24 hour battery life -11.158 Over-ear headphones with deep bass and a foldable design -11.278 Budget laptop for students, lightweight with long battery life
Both agree on the top pick. Look further down: the bi-encoder ranked "wireless earbuds with noise cancellation" fourth, behind an unrelated laptop. The cross-encoder, reading query and document together, correctly recognises the earbuds as more relevant than the laptop and moves it up to third, pushing the laptop down to last.
The walkthrough
The cross-encoder's numbers are not probabilities or similarities. They are raw, un-normalised relevance logits from cross-encoder/ms-marco-MiniLM-L-6-v2, trained on the MS MARCO passage-ranking dataset. Negative numbers are completely normal — only the relative order matters, not the absolute value.
CrossEncoder.predict() takes (query, document) pairs, not separate lists. This is the API's way of enforcing the whole point of a cross-encoder: query and document are concatenated and passed through the model together, as a single input, not encoded independently and compared afterward.
Why only 5 candidates were re-scored, not all 15. This mirrors real production usage: a cross-encoder is too slow to run against an entire catalogue for every search, so it only ever re-judges a short list a faster method already narrowed down. The full two-stage pattern gets its own lesson next.
Common mistakes
Using a cross-encoder as your only retrieval method. It has no way to search a large catalogue efficiently — there is no such thing as "precompute the cross-encoder score" for a document, since the score does not exist until the query is known.
Treating cross-encoder scores as comparable across different queries. A score of -3.9 for one query and -3.9 for a completely different query say nothing about how those two queries compare to each other. Only rank within a single query's candidate list.
Forgetting bi-encoder vectors can be precomputed and reused. The single biggest performance win in this whole comparison is that doc_vecs only ever needs to be computed once, no matter how many searches follow. Recomputing it per query throws away the entire speed advantage a bi-encoder offers.
Try it yourself
Swap the model to the smaller cross-encoder/ms-marco-TinyBERT-L-2-v2 (a lighter, faster cross-encoder, a few times smaller) and compare its ranking against the one above. Smaller cross-encoders trade some accuracy for speed, the same trade-off shape as bi-encoder versus cross-encoder itself, one level down.
What to learn next
- Reranking the top 50 — the full two-stage pipeline this lesson previewed.
- ColBERT and late interaction — a third architecture between bi-encoders and cross-encoders.
- Attention — the mechanism giving a cross-encoder its query-document interaction.
Researcher — Mathematics and papers.
Formal comparison
A bi-encoder computes:
score(q, d) = sim( f(q), f(d) )Where f is a shared encoder applied independently to query q and document d, and sim is a cheap similarity function (cosine similarity, or a dot product for normalised vectors). Because f(d) never depends on q, every document embedding can be computed once, offline, and reused for every future query.
A cross-encoder computes:
score(q, d) = g([q ; d])Where g is a single transformer applied to the concatenation of query and document (typically [CLS] query [SEP] document [SEP]), with full self-attention between every query token and every document token, and a scalar relevance score read from the [CLS] output. Nothing about g([q;d]) can be precomputed independently of q, because attention couples every token of d to every token of q from the first layer onward.
Why cross-encoders are more accurate
The attention mechanism lets a cross-encoder model fine-grained token-level interactions — negation, numeric comparison, exact entity matching — that a fixed-size pooled vector from a bi-encoder can dilute or lose entirely. Nogueira & Cho (2019), Passage Re-ranking with BERT, showed a BERT cross-encoder (monoBERT) reranking BM25's initial results produced a substantial jump in MS MARCO passage-ranking metrics over bi-encoder or BM25 baselines alone, establishing the retrieve-then-rerank pattern as a de facto standard.
The cost, quantified
Reimers & Gurevych's original Sentence-BERT paper (2019) makes the scaling argument concretely: finding the most similar pair among 10,000 sentences with a BERT cross-encoder requires roughly C(10,000, 2) ≈ 50 million inference passes (all pairs of 10,000 sentences), reported in the paper as on the order of 65 hours on a V100 GPU — versus roughly 5 seconds to embed all 10,000 sentences once with a bi-encoder and compute pairwise cosine similarities. That gap is the entire reason bi-encoders exist as a distinct architecture rather than everyone using cross-encoders everywhere.
Key references
- Reimers, N. & Gurevych, I. (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. arXiv:1908.10084
- Nogueira, R. & Cho, K. (2019). Passage Re-ranking with BERT. arXiv:1901.04085
- Humeau, S., Shuster, K., Lachaux, M-A. & Weston, J. (2020). Poly-encoders: Transformer Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence Scoring. ICLR.
Current state and middle-ground architectures
Poly-encoders (Humeau et al., 2020) sit between the two: they compute several context vectors per document (not only one, unlike a standard bi-encoder), attended over by the query at scoring time, recovering some cross-attention benefit while keeping documents pre-encodable. Late-interaction models — covered later in this section — push further in this direction, keeping a full set of per-token document vectors rather than a single pooled one. The retrieve-with-a-bi-encoder, rerank-with-a-cross-encoder pattern from the developer block remains the dominant production architecture as of the current generation of search systems, precisely because it lets each stage do the job it is actually efficient at.
What to learn next
- Reranking the top 50 — the full two-stage pipeline this lesson previewed.
- ColBERT and late interaction — a third architecture between bi-encoders and cross-encoders.
- Attention — the mechanism giving a cross-encoder its query-document interaction.