Semantic Search and Reranking

Hybrid search with reciprocal rank fusion

Combine a keyword ranking and a semantic ranking into one list using nothing but their rank positions — a simple formula that is remarkably hard to beat.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Hybrid search asks two different search methods for their opinion, then merges the two rankings into one.

Ask two friends where to eat. One always remembers exact restaurant names and dishes. Ask for "that biryani place near the station," and they nail it instantly. The other doesn't remember names at all, but always understands what you're actually craving, even from a vague description. On any given night, you would rather have both opinions than pick one friend and ignore the other.

Hybrid search does exactly this with keyword search and semantic search. Neither one is reliably better across every query. Keyword search nails exact terms. Semantic search nails paraphrased meaning. Hybrid search runs both, and merges their two ranked lists into one sturdier list.

Why it exists

The previous lessons showed keyword and semantic search each failing in situations the other handles fine. Picking one search method means permanently accepting its blind spots. Running both, and merging their results, covers each one's weak spot with the other's strength. That only works if the merging step is done properly.

How it works

query: "something to bring peace and quiet on a long flight"

BM25 ranking (by position)         Semantic ranking (by position)
 1. wireless mouse                  1. flight headphones  <-- the right answer
 2. budget laptop                   2. bluetooth speaker
 3. flight headphones                3. budget laptop
 4. ...                              4. ...

RECIPROCAL RANK FUSION: score each document using ONLY its rank
position in each list, not the raw scores (which aren't comparable
between BM25 and cosine similarity anyway)

  score(doc) = 1/(60 + rank in BM25 list) + 1/(60 + rank in semantic list)

fused ranking:
 1. flight headphones   <- BM25 said 3rd, semantic said 1st -- still wins overall
 2. budget laptop
 3. ...

The trick is fusing by rank position, not raw score. BM25's numbers and cosine similarities live on completely different, incomparable scales. Combining them directly would be meaningless.

A real example you have seen

Many modern search boxes handle both an exact product name and a vague, meandering description well. Some even handle the vague one better. That usually means a keyword signal and a semantic signal are being blended behind the scenes. Neither weak spot gets exposed to the user.

Remember this

  • Hybrid search combines a keyword ranking and a semantic ranking rather than betting on either alone.
  • Combine by rank position, not raw score — BM25 scores and cosine similarities are not on comparable scales.
  • A document weak in one ranking can still win overall if it is strong in the other.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install rank_bm25 sentence-transformers

Outputs verified with rank_bm25 0.2.2 and sentence-transformers 5.4.1 on CPU.

Fusing BM25 and semantic rankings with Reciprocal Rank Fusion

hybrid_rrf.py
import numpy as np
from rank_bm25 import BM25Okapi
from sentence_transformers import SentenceTransformer

products = [
    "Wireless earbuds with noise cancellation and 24 hour battery life",
    "Over-ear headphones with deep bass and a foldable design",
    "Bluetooth speaker, waterproof and portable, great for outdoor parties",
    "Laptop with 16GB RAM and fast SSD storage, built for programming",
    "Budget laptop for students, lightweight with long battery life",
    "Mechanical keyboard with RGB lighting and tactile switches",
    "Wireless mouse with an ergonomic shape that works on any surface",
    "Smartphone with a triple camera and all-day battery",
    "Budget smartphone with a big screen for watching videos",
    "Smartwatch that tracks your heart rate and sleep",
    "Fitness band with a step counter and water resistance",
    "External hard drive with 1TB of storage and fast USB-C transfer",
    "Portable power bank that charges your phone twice over",
    "Over-ear headphones built for long flights, with active noise cancelling",
    "4K webcam for video calls and live streaming",
]
query = "something to bring peace and quiet on a long flight"

bm25 = BM25Okapi([p.lower().split() for p in products])
bm25_scores = bm25.get_scores(query.lower().split())
bm25_ranks = {doc_id: rank for rank, doc_id in enumerate(np.argsort(bm25_scores)[::-1], start=1)}

model = SentenceTransformer("all-MiniLM-L6-v2")
doc_vecs = model.encode(products, normalize_embeddings=True)
q_vec = model.encode([query], normalize_embeddings=True)[0]
sem_scores = doc_vecs @ q_vec
sem_ranks = {doc_id: rank for rank, doc_id in enumerate(np.argsort(sem_scores)[::-1], start=1)}


def reciprocal_rank_fusion(rank_lists, k=60):
    fused = {}
    for ranks in rank_lists:
        for doc_id, rank in ranks.items():
            fused[doc_id] = fused.get(doc_id, 0.0) + 1.0 / (k + rank)
    return fused


fused_scores = reciprocal_rank_fusion([bm25_ranks, sem_ranks])
fused_order = sorted(fused_scores, key=fused_scores.get, reverse=True)

print(f"{'rank':>4}  {'BM25 rank':>9}  {'semantic rank':>13}  {'RRF score':>9}  product")
for rank, doc_id in enumerate(fused_order[:5], start=1):
    print(f"{rank:>4}  {bm25_ranks[doc_id]:>9}  {sem_ranks[doc_id]:>13}  {fused_scores[doc_id]:>9.4f}  {products[doc_id]}")
Output
rank  BM25 rank  semantic rank  RRF score  product
   1          3              1     0.0323  Over-ear headphones built for long flights, with active noise cancelling
   2          2              3     0.0320  Budget laptop for students, lightweight with long battery life
   3          4              7     0.0306  Smartphone with a triple camera and all-day battery
   4          6              5     0.0305  Over-ear headphones with deep bass and a foldable design
   5         11              2     0.0302  Bluetooth speaker, waterproof and portable, great for outdoor parties

The right answer — flight headphones — sat third in BM25's ranking, not first. Because it was strongly first-place in the semantic ranking, the fused score still puts it at the top, robust to BM25's mistake here rather than inheriting it.

The walkthrough

Ranks are built with enumerate(..., start=1) on the sorted order, not the raw scores themselves. This is the step that makes fusion possible at all — a rank of "3rd" means the same thing whether it came from a BM25 score of 2.09 or a cosine similarity of 0.42, which is precisely why fusing by rank sidesteps the incomparable-scales problem from the previous lesson.

k=60 in the formula is a smoothing constant, not a candidate count. It softens the impact of small rank differences near the top of each list — the difference between rank 1 and rank 2 matters less to the final score than it would with k=0, which keeps one list's noisy top pick from single-handedly dominating the fused result.

A document missing entirely from one of the two ranked lists still gets scored. fused.get(doc_id, 0.0) treats an absent rank contribution as zero rather than raising an error — useful when the two retrieval methods return different-sized candidate sets in a real system.

Common mistakes

Averaging raw BM25 scores and cosine similarities directly. These live on entirely different, corpus-dependent scales. A BM25 score of 2.0 and a cosine similarity of 2.0 mean nothing comparable — combining them without first converting to ranks (or otherwise normalising) produces a fused list dominated by whichever method happens to produce numerically larger raw scores.

Assuming hybrid search always beats each individual method. On this exact query, the fused top pick matches what semantic search alone already found — fusion did not need to do much work here. Its real value shows up in aggregate, across many queries with different keyword-versus-semantic strengths, not necessarily on any single example.

Picking k=60 without knowing it is a convention, not a law. The original Reciprocal Rank Fusion paper used k=60 and it has become a common default, but it is a tunable hyperparameter — smaller values weight the top of each list more heavily.

Try it yourself

Try the query "RGB mechanical keyboard" from the keyword-vs-semantic lesson, where both BM25 and semantic search already agreed on the top result. Confirm the fused ranking agrees too — fusion should not disturb a case where both underlying methods already point the same way.

What to learn next

Researcher — Mathematics and papers.

The formula

Given a set of ranked lists R_1, ..., R_m (one per retrieval system) over a common candidate pool, Reciprocal Rank Fusion scores each document d as:

text
RRF(d) = sum over i=1..m of  1 / (k + rank_i(d))

Where rank_i(d) is d's 1-indexed position in ranked list i (treated as infinity, contributing 0, if d does not appear in that list), and k is a constant, k=60 in the original paper.

Why rank-based fusion, and why it works well empirically

Cormack, Clarke & Buettcher (2009), Reciprocal Rank Fusion outperforms Condorcet and individual Rank Learning Methods, introduced RRF as a deliberately score-agnostic alternative to fusion methods that require normalising and combining raw relevance scores across systems with incompatible scales and distributions — exactly the BM25-versus-cosine-similarity problem this lesson demonstrates directly. Because RRF only consumes rank positions, it needs no score calibration between systems, no knowledge of each system's internals, and extends readily to combining more than two ranked lists.

The paper's central empirical finding, across TREC ranking benchmarks, was that simple reciprocal-rank averaging outperformed considerably more sophisticated rank-aggregation methods (Condorcet fusion, learned rank-combination models) on the tasks tested — a result the authors themselves note is somewhat surprising given the method's simplicity, and part of why RRF remains the default fusion choice in production hybrid search systems rather than more elaborate learned alternatives.

Alternatives to RRF

Convex combination of normalised scores — min-max or z-score normalising each system's raw scores before a weighted sum — can outperform RRF when the score distributions are well-behaved and the relative weighting between systems is tuned, but is more fragile: a shift in one system's score distribution (a model update, a different query type) silently changes the effective weighting.

Learned fusion (learning-to-rank) treats the outputs of each retrieval system as input features to a supervised ranking model (commonly gradient-boosted trees, following the LambdaMART line of work), trained on relevance-labelled query-document pairs. This can outperform RRF given enough labelled training data, at the cost of needing that data and a retraining pipeline RRF does not require.

Key references

  • Cormack, G., Clarke, C. & Buettcher, S. (2009). Reciprocal Rank Fusion outperforms Condorcet and individual Rank Learning Methods. SIGIR.
  • Bruch, S., Gai, S. & Ingber, A. (2023). An Analysis of Fusion Functions for Hybrid Retrieval. arXiv:2210.11934

Current state

RRF is the default hybrid-search fusion method offered out of the box by most current vector database and search platforms, precisely because it requires no per-deployment score calibration. Bruch, Gai & Ingber (2023) provide a more recent, systematic comparison of fusion functions across retrieval settings, generally confirming RRF as a strong, low-maintenance default while identifying conditions — well-calibrated score distributions, sufficient tuning budget — under which weighted score combination can edge it out.

What to learn next