Hybrid search with reciprocal rank fusion
Combine a keyword ranking and a semantic ranking into one list using nothing but their rank positions — a simple formula that is remarkably hard to beat.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Hybrid search asks two different search methods for their opinion, then merges the two rankings into one.
Ask two friends where to eat. One always remembers exact restaurant names and dishes. Ask for "that biryani place near the station," and they nail it instantly. The other doesn't remember names at all, but always understands what you're actually craving, even from a vague description. On any given night, you would rather have both opinions than pick one friend and ignore the other.
Hybrid search does exactly this with keyword search and semantic search. Neither one is reliably better across every query. Keyword search nails exact terms. Semantic search nails paraphrased meaning. Hybrid search runs both, and merges their two ranked lists into one sturdier list.
Why it exists
The previous lessons showed keyword and semantic search each failing in situations the other handles fine. Picking one search method means permanently accepting its blind spots. Running both, and merging their results, covers each one's weak spot with the other's strength. That only works if the merging step is done properly.
How it works
query: "something to bring peace and quiet on a long flight"
BM25 ranking (by position) Semantic ranking (by position)
1. wireless mouse 1. flight headphones <-- the right answer
2. budget laptop 2. bluetooth speaker
3. flight headphones 3. budget laptop
4. ... 4. ...
RECIPROCAL RANK FUSION: score each document using ONLY its rank
position in each list, not the raw scores (which aren't comparable
between BM25 and cosine similarity anyway)
score(doc) = 1/(60 + rank in BM25 list) + 1/(60 + rank in semantic list)
fused ranking:
1. flight headphones <- BM25 said 3rd, semantic said 1st -- still wins overall
2. budget laptop
3. ...The trick is fusing by rank position, not raw score. BM25's numbers and cosine similarities live on completely different, incomparable scales. Combining them directly would be meaningless.
A real example you have seen
Many modern search boxes handle both an exact product name and a vague, meandering description well. Some even handle the vague one better. That usually means a keyword signal and a semantic signal are being blended behind the scenes. Neither weak spot gets exposed to the user.
Remember this
- Hybrid search combines a keyword ranking and a semantic ranking rather than betting on either alone.
- Combine by rank position, not raw score — BM25 scores and cosine similarities are not on comparable scales.
- A document weak in one ranking can still win overall if it is strong in the other.
What to learn next
- How HNSW and IVF actually work — making the semantic half of a hybrid pipeline fast at real scale.
- FAISS — a concrete library for the vector-search half of hybrid retrieval.
- Building a retrieval test set for your own corpus — measuring whether hybrid search is actually an improvement on your data.
Developer — Code and libraries.
Setup
pip install rank_bm25 sentence-transformersOutputs verified with rank_bm25 0.2.2 and sentence-transformers 5.4.1 on CPU.
Fusing BM25 and semantic rankings with Reciprocal Rank Fusion
import numpy as np
from rank_bm25 import BM25Okapi
from sentence_transformers import SentenceTransformer
products = [
"Wireless earbuds with noise cancellation and 24 hour battery life",
"Over-ear headphones with deep bass and a foldable design",
"Bluetooth speaker, waterproof and portable, great for outdoor parties",
"Laptop with 16GB RAM and fast SSD storage, built for programming",
"Budget laptop for students, lightweight with long battery life",
"Mechanical keyboard with RGB lighting and tactile switches",
"Wireless mouse with an ergonomic shape that works on any surface",
"Smartphone with a triple camera and all-day battery",
"Budget smartphone with a big screen for watching videos",
"Smartwatch that tracks your heart rate and sleep",
"Fitness band with a step counter and water resistance",
"External hard drive with 1TB of storage and fast USB-C transfer",
"Portable power bank that charges your phone twice over",
"Over-ear headphones built for long flights, with active noise cancelling",
"4K webcam for video calls and live streaming",
]
query = "something to bring peace and quiet on a long flight"
bm25 = BM25Okapi([p.lower().split() for p in products])
bm25_scores = bm25.get_scores(query.lower().split())
bm25_ranks = {doc_id: rank for rank, doc_id in enumerate(np.argsort(bm25_scores)[::-1], start=1)}
model = SentenceTransformer("all-MiniLM-L6-v2")
doc_vecs = model.encode(products, normalize_embeddings=True)
q_vec = model.encode([query], normalize_embeddings=True)[0]
sem_scores = doc_vecs @ q_vec
sem_ranks = {doc_id: rank for rank, doc_id in enumerate(np.argsort(sem_scores)[::-1], start=1)}
def reciprocal_rank_fusion(rank_lists, k=60):
fused = {}
for ranks in rank_lists:
for doc_id, rank in ranks.items():
fused[doc_id] = fused.get(doc_id, 0.0) + 1.0 / (k + rank)
return fused
fused_scores = reciprocal_rank_fusion([bm25_ranks, sem_ranks])
fused_order = sorted(fused_scores, key=fused_scores.get, reverse=True)
print(f"{'rank':>4} {'BM25 rank':>9} {'semantic rank':>13} {'RRF score':>9} product")
for rank, doc_id in enumerate(fused_order[:5], start=1):
print(f"{rank:>4} {bm25_ranks[doc_id]:>9} {sem_ranks[doc_id]:>13} {fused_scores[doc_id]:>9.4f} {products[doc_id]}")rank BM25 rank semantic rank RRF score product 1 3 1 0.0323 Over-ear headphones built for long flights, with active noise cancelling 2 2 3 0.0320 Budget laptop for students, lightweight with long battery life 3 4 7 0.0306 Smartphone with a triple camera and all-day battery 4 6 5 0.0305 Over-ear headphones with deep bass and a foldable design 5 11 2 0.0302 Bluetooth speaker, waterproof and portable, great for outdoor parties
The right answer — flight headphones — sat third in BM25's ranking, not first. Because it was strongly first-place in the semantic ranking, the fused score still puts it at the top, robust to BM25's mistake here rather than inheriting it.
The walkthrough
Ranks are built with enumerate(..., start=1) on the sorted order, not the raw scores themselves. This is the step that makes fusion possible at all — a rank of "3rd" means the same thing whether it came from a BM25 score of 2.09 or a cosine similarity of 0.42, which is precisely why fusing by rank sidesteps the incomparable-scales problem from the previous lesson.
k=60 in the formula is a smoothing constant, not a candidate count. It softens the impact of small rank differences near the top of each list — the difference between rank 1 and rank 2 matters less to the final score than it would with k=0, which keeps one list's noisy top pick from single-handedly dominating the fused result.
A document missing entirely from one of the two ranked lists still gets scored. fused.get(doc_id, 0.0) treats an absent rank contribution as zero rather than raising an error — useful when the two retrieval methods return different-sized candidate sets in a real system.
Common mistakes
Averaging raw BM25 scores and cosine similarities directly. These live on entirely different, corpus-dependent scales. A BM25 score of 2.0 and a cosine similarity of 2.0 mean nothing comparable — combining them without first converting to ranks (or otherwise normalising) produces a fused list dominated by whichever method happens to produce numerically larger raw scores.
Assuming hybrid search always beats each individual method. On this exact query, the fused top pick matches what semantic search alone already found — fusion did not need to do much work here. Its real value shows up in aggregate, across many queries with different keyword-versus-semantic strengths, not necessarily on any single example.
Picking k=60 without knowing it is a convention, not a law. The original Reciprocal Rank Fusion paper used k=60 and it has become a common default, but it is a tunable hyperparameter — smaller values weight the top of each list more heavily.
Try it yourself
Try the query "RGB mechanical keyboard" from the keyword-vs-semantic lesson, where both BM25 and semantic search already agreed on the top result. Confirm the fused ranking agrees too — fusion should not disturb a case where both underlying methods already point the same way.
What to learn next
- How HNSW and IVF actually work — making the semantic half of a hybrid pipeline fast at real scale.
- FAISS — a concrete library for the vector-search half of hybrid retrieval.
- Building a retrieval test set for your own corpus — measuring whether hybrid search is actually an improvement on your data.
Researcher — Mathematics and papers.
The formula
Given a set of ranked lists R_1, ..., R_m (one per retrieval system) over a common candidate pool, Reciprocal Rank Fusion scores each document d as:
RRF(d) = sum over i=1..m of 1 / (k + rank_i(d))Where rank_i(d) is d's 1-indexed position in ranked list i (treated as infinity, contributing 0, if d does not appear in that list), and k is a constant, k=60 in the original paper.
Why rank-based fusion, and why it works well empirically
Cormack, Clarke & Buettcher (2009), Reciprocal Rank Fusion outperforms Condorcet and individual Rank Learning Methods, introduced RRF as a deliberately score-agnostic alternative to fusion methods that require normalising and combining raw relevance scores across systems with incompatible scales and distributions — exactly the BM25-versus-cosine-similarity problem this lesson demonstrates directly. Because RRF only consumes rank positions, it needs no score calibration between systems, no knowledge of each system's internals, and extends readily to combining more than two ranked lists.
The paper's central empirical finding, across TREC ranking benchmarks, was that simple reciprocal-rank averaging outperformed considerably more sophisticated rank-aggregation methods (Condorcet fusion, learned rank-combination models) on the tasks tested — a result the authors themselves note is somewhat surprising given the method's simplicity, and part of why RRF remains the default fusion choice in production hybrid search systems rather than more elaborate learned alternatives.
Alternatives to RRF
Convex combination of normalised scores — min-max or z-score normalising each system's raw scores before a weighted sum — can outperform RRF when the score distributions are well-behaved and the relative weighting between systems is tuned, but is more fragile: a shift in one system's score distribution (a model update, a different query type) silently changes the effective weighting.
Learned fusion (learning-to-rank) treats the outputs of each retrieval system as input features to a supervised ranking model (commonly gradient-boosted trees, following the LambdaMART line of work), trained on relevance-labelled query-document pairs. This can outperform RRF given enough labelled training data, at the cost of needing that data and a retraining pipeline RRF does not require.
Key references
- Cormack, G., Clarke, C. & Buettcher, S. (2009). Reciprocal Rank Fusion outperforms Condorcet and individual Rank Learning Methods. SIGIR.
- Bruch, S., Gai, S. & Ingber, A. (2023). An Analysis of Fusion Functions for Hybrid Retrieval. arXiv:2210.11934
Current state
RRF is the default hybrid-search fusion method offered out of the box by most current vector database and search platforms, precisely because it requires no per-deployment score calibration. Bruch, Gai & Ingber (2023) provide a more recent, systematic comparison of fusion functions across retrieval settings, generally confirming RRF as a strong, low-maintenance default while identifying conditions — well-calibrated score distributions, sufficient tuning budget — under which weighted score combination can edge it out.
What to learn next
- How HNSW and IVF actually work — making the semantic half of a hybrid pipeline fast at real scale.
- FAISS — a concrete library for the vector-search half of hybrid retrieval.
- Building a retrieval test set for your own corpus — measuring whether hybrid search is actually an improvement on your data.