Semantic Search and Reranking

Building a retrieval test set for your own corpus

Without a set of queries and known-correct answers, there is no way to tell whether a search change actually helped. This lesson builds one and computes the two metrics that matter most.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A retrieval eval set is a list of questions with the correct answers already written down. It lets you score a search system honestly, instead of guessing.

A teacher grading an exam needs an answer key made before grading starts. Otherwise "grading" is only reading answers and deciding, in the moment, whether each one feels right. An answer key fixes what "correct" means in advance. Every answer sheet then gets judged by the same standard. Two different teachers would grade it the same way.

A retrieval eval set is that answer key for search. It is a list of realistic queries, each paired with the documents a person has already judged correct. Hybrid search, reranking, query expansion — every technique in this section needs a fixed answer key like this one. Only then can it be judged "better" or "worse," honestly.

Why it exists

Every earlier lesson made a claim like "this retrieved the right document higher up," using one example query. That query was chosen because it demonstrated the point well. That is fine for teaching. It is not how you decide whether a change to a real search system actually helped. One hand-picked example proves nothing about the other thousand queries a real system handles.

An eval set turns "I think this change helped" into a number you can compare, before and after. It runs across a representative set of queries, chosen ahead of time, never cherry-picked afterward to flatter the result.

How it works

build the eval set FIRST, before testing anything:

  query: "gift for someone who exercises"
  relevant documents: {smartwatch, fitness band}     <- decided by a human, in advance

  query: "affordable laptop for college"
  relevant documents: {budget laptop}

  ... a handful more, covering different query styles

run your search system against every query in the set
        |
        v
for each query, check:
  - did the correct document(s) appear in the top few results at all?  (RECALL)
  - how far down the list was the FIRST correct one?                   (RANK)
        |
        v
average both across every query in the set -- now you have one honest,
comparable number for "how good is this search system today"

The eval set only needs to be built once. After that, every future change to the search pipeline gets checked against the exact same fixed standard.

A real example you have seen

A search team that says "our model improved results by 12%" is quoting a number computed exactly this way. A fixed set of queries, with known-correct answers, gets scored before and after the change. That scoring runs on the same held-back test set neither model was tuned against.

Remember this

  • An eval set is a fixed list of queries with pre-judged correct answers, built before you start testing changes.
  • Recall@k asks whether the correct answer showed up at all, within the top k results.
  • MRR (mean reciprocal rank) asks how far down the list the first correct answer landed — rewarding it appearing early, not only appearing.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install sentence-transformers

Outputs verified with sentence-transformers 5.4.1 on CPU.

A five-query eval set, scored with Recall@3 and MRR

retrieval_eval.py
import numpy as np
from sentence_transformers import SentenceTransformer

products = [
    "Wireless earbuds with noise cancellation and 24 hour battery life",
    "Over-ear headphones with deep bass and a foldable design",
    "Bluetooth speaker, waterproof and portable, great for outdoor parties",
    "Laptop with 16GB RAM and fast SSD storage, built for programming",
    "Budget laptop for students, lightweight with long battery life",
    "Mechanical keyboard with RGB lighting and tactile switches",
    "Wireless mouse with an ergonomic shape that works on any surface",
    "Smartphone with a triple camera and all-day battery",
    "Budget smartphone with a big screen for watching videos",
    "Smartwatch that tracks your heart rate and sleep",
    "Fitness band with a step counter and water resistance",
    "External hard drive with 1TB of storage and fast USB-C transfer",
    "Portable power bank that charges your phone twice over",
    "Over-ear headphones built for long flights, with active noise cancelling",
    "4K webcam for video calls and live streaming",
]
model = SentenceTransformer("all-MiniLM-L6-v2")
doc_vecs = model.encode(products, normalize_embeddings=True)

# the eval set: each query paired with the product ids a human judged relevant
EVAL_SET = [
    ("something to bring peace and quiet on a long flight", {0, 13}),
    ("gift for someone who exercises", {9, 10}),
    ("device for taking video calls from home", {14}),
    ("affordable laptop for college", {4}),
    ("keyboard and mouse for a new desk setup", {5, 6}),
]
K = 3


def search(q, k=K):
    q_vec = model.encode([q], normalize_embeddings=True)[0]
    scores = doc_vecs @ q_vec
    return list(np.argsort(scores)[::-1][:k])


recalls, reciprocal_ranks = [], []
for query, relevant in EVAL_SET:
    ranked = search(query, k=len(products))
    top_k = set(ranked[:K])
    recall = len(top_k & relevant) / len(relevant)
    recalls.append(recall)

    rr = 0.0
    for rank, doc_id in enumerate(ranked, start=1):
        if doc_id in relevant:
            rr = 1.0 / rank
            break
    reciprocal_ranks.append(rr)

    print(f"query: {query!r}")
    print(f"  relevant ids: {sorted(relevant)}   top-{K} retrieved: {ranked[:K]}")
    print(f"  recall@{K}: {recall:.2f}   reciprocal rank: {rr:.2f}")

print(f"\nmean recall@{K}: {np.mean(recalls):.2f}")
print(f"MRR: {np.mean(reciprocal_ranks):.2f}")
Output
query: 'something to bring peace and quiet on a long flight'
  relevant ids: [0, 13]   top-3 retrieved: [13, 2, 4]
  recall@3: 0.50   reciprocal rank: 1.00
query: 'gift for someone who exercises'
  relevant ids: [9, 10]   top-3 retrieved: [10, 2, 4]
  recall@3: 0.50   reciprocal rank: 1.00
query: 'device for taking video calls from home'
  relevant ids: [14]   top-3 retrieved: [8, 14, 7]
  recall@3: 1.00   reciprocal rank: 0.50
query: 'affordable laptop for college'
  relevant ids: [4]   top-3 retrieved: [4, 3, 8]
  recall@3: 1.00   reciprocal rank: 1.00
query: 'keyboard and mouse for a new desk setup'
  relevant ids: [5, 6]   top-3 retrieved: [5, 6, 3]
  recall@3: 1.00   reciprocal rank: 1.00

mean recall@3: 0.80
MRR: 0.90

A genuinely mixed result, and that is what an honest eval run looks like. Two queries only found one of their two relevant items within the top 3 (recall 0.50); one query found its single relevant item, but only at rank 2 (reciprocal rank 0.50); two queries scored perfectly on both metrics. None of the earlier lessons' single hand-picked examples showed this kind of nuance — that is exactly why a proper eval set, run across several queries at once, matters.

The walkthrough

The eval set is written by hand, by someone using judgment, before any scoring happens. EVAL_SET encodes five genuine human relevance judgments. This step cannot be automated away — it is the one piece of ground truth every metric downstream depends on entirely.

Recall@k asks a yes-or-no-per-item question: did each relevant document appear at all, within the top k? len(top_k & relevant) / len(relevant) is the fraction of relevant items that were successfully retrieved — it does not care about their exact position within the top k, only whether they made the cut at all.

MRR asks a different question: how early did the first relevant result appear? For the video-call query, the webcam was genuinely retrieved — but at rank 2, not rank 1, giving a reciprocal rank of 0.5 rather than a full 1.0. Recall@3 for that query is a perfect 1.00 despite this, because recall only checks whether it appeared within the top 3 at all — the two metrics are answering genuinely different questions, which is why both are reported.

Averaging across all five queries produces the two headline numbers. mean recall@3: 0.80 and MRR: 0.90 are what you would actually report and track over time — not any single query's individual score.

Common mistakes

Building the eval set from the same queries used to tune the system. This overstates real-world performance, the same way testing a student only on questions from their own practice sheet would. A proper eval set represents queries the system was not specifically tuned against.

Judging relevance loosely, "at a glance," instead of deciding it in advance and writing it down. The whole value of an eval set comes from its judgments being fixed before scoring — deciding relevance on the fly, while looking at results, quietly turns the eval set into confirmation of whatever the system already did.

Reporting only one metric. As the video-call query shows directly, recall@3 and MRR can tell different stories about the exact same result. A dashboard tracking only one of them can hide a real regression the other would have caught.

Drawing firm conclusions from five queries. This eval set is intentionally tiny, for a runnable demonstration. A real eval set typically needs dozens to hundreds of queries, spanning different query styles and difficulty levels, before its aggregate numbers are trustworthy enough to make a shipping decision on.

Try it yourself

Add a sixth query where you deliberately expect the current search to fail — something vague or oddly worded, in the spirit of the query expansion lesson — and confirm its recall and reciprocal rank actually come out low. A good eval set should contain queries the current system is known to struggle with, not only ones it already handles well.

What to learn next

Researcher — Mathematics and papers.

Metric definitions, formally

For a query q with relevant document set R_q and a ranked result list of length k, D_q(k):

text
Recall@k(q) = |R_q ∩ D_q(k)| / |R_q|

RR(q) = 1 / (rank of first relevant result)
MRR   = (1 / |Q|) * sum over q in Q of RR(q)

MRR, introduced into standard IR evaluation practice via the TREC Question Answering track (Voorhees, 1999), is most appropriate when a query typically has one correct answer and early position matters most — as in the video-call query above. Recall@k is more appropriate when a query can have multiple valid relevant documents and the goal is coverage within a result page, not strictly the very first position.

Beyond binary relevance: nDCG

Both metrics above treat relevance as binary — a document either counts as relevant or it doesn't. Normalised Discounted Cumulative Gain (Järvelin & Kekäläinen, 2002) generalises to graded relevance (a document can be "somewhat relevant," "very relevant," and so on) and applies a logarithmic discount to reward relevant documents appearing earlier without an all-or-nothing cutoff at position k:

text
DCG@k  = sum over i=1..k of (2^rel_i - 1) / log2(i + 1)
nDCG@k = DCG@k / IDCG@k

Where rel_i is the graded relevance of the document at rank i, and IDCG@k is the DCG@k of the ideal (perfectly sorted) ranking, used to normalise the score into the range 0 to 1. nDCG is the standard metric in large-scale IR benchmarks precisely because most real search systems care about more than a binary relevant/not-relevant judgment.

Standardised benchmarks

Hand-built eval sets like the one in the developer block are essential for a specific corpus and use case, but the field also relies on standardised, reusable benchmarks for comparing general-purpose retrieval methods. BEIR (Thakur, Reimers, Rücklé, Srivastava & Gurevych, 2021) assembles 18 diverse retrieval datasets across domains (biomedical, financial, question answering, fact-checking, and more) specifically to test zero-shot generalisation — whether a retrieval method trained on one distribution of queries and documents holds up on a genuinely different one, which single-domain eval sets like the toy example above cannot measure at all.

Key references

  • Voorhees, E. (1999). The TREC-8 Question Answering Track Report. TREC.
  • Järvelin, K. & Kekäläinen, J. (2002). Cumulated Gain-Based Evaluation of IR Techniques. ACM Transactions on Information Systems 20(4).
  • Thakur, N., Reimers, N., Rücklé, A., Srivastava, A. & Gurevych, I. (2021). BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models. arXiv:2104.08663

Current state

Recall@k and MRR remain the two most commonly reported metrics for practical, corpus-specific retrieval evaluation because they are cheap to compute and easy to explain to a non-specialist stakeholder; nDCG is standard wherever graded relevance judgments are available and the full ranking, not only the first hit, genuinely matters. For any team building or changing a real search system, the single highest-leverage step described across this entire section is the one covered in this lesson: build the eval set before making any changes at all, so every later claim of improvement has something honest to be measured against.

What to learn next