Semantic Search and Reranking

Keyword search vs semantic search

Keyword search matches the exact words you typed. Semantic search matches what you meant. The difference shows up the moment your words don't match the document's words.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Keyword search finds documents containing your exact words. Semantic search finds documents that mean what you meant, even with completely different words.

Walk into a small shop and ask for "something to clean my hands." A good shopkeeper walks you to the hand soap without blinking. You never even said the word "soap." A shopkeeper who only reacts to the exact word "soap" would stare at you blankly. Only saying "soap" out loud would work.

Search engines have both kinds of shopkeeper built into them. Keyword search is the second one. It matches the literal words in your query against the literal words in each document. Semantic search is the first. It compares what your query means against what each document means, using embeddings — numeric representations of meaning.

Why it exists

For decades, search ran on keyword matching. It was refined into a genuinely useful formula called BM25. It scores documents by how often, and how distinctively, they contain your exact query words. It works well whenever people search using the same words the documents use.

The problem shows up constantly in practice. You search "affordable laptop," and the product listing says "budget-friendly notebook." Zero shared words, same meaning. Keyword search misses it. This gap is exactly what semantic search was built to close, by comparing meaning instead of spelling.

How it works

query: "something to bring peace and quiet on a long flight"

KEYWORD SEARCH (BM25)                    SEMANTIC SEARCH (embeddings)
checks: which exact words match?         checks: which MEANING is closest?

  "wireless mouse ... any surface"          "over-ear headphones ... for
   matches "a" -- ranked HIGH               long flights, noise cancelling"
   (barely relevant)                        ranked HIGHEST (genuinely relevant,
                                             despite sharing almost no exact words
  the actually-relevant headphones          with the query)
   product ranks LOWER, because it
   shares fewer literal words with
   the query than you'd expect

Keyword search counts overlapping words. Semantic search measures distance between two points in a learned "meaning space." The query and the right document can land close together there, even without sharing vocabulary.

A real example you have seen

Search your phone's photo gallery for "dog" and it finds pictures never captioned "dog" at all. The app recognised what was in the photo. A vague, oddly-worded question into a modern search bar can still land on the right page. That is the text version of the same shift: understanding intent, not only spelling.

Remember this

  • Keyword search matches literal words; semantic search matches meaning.
  • Keyword search fails the moment your words differ from the document's words, even when the meaning is identical.
  • Neither approach is strictly better — hybrid search, later in this section, combines both.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install rank_bm25 sentence-transformers

rank_bm25 is a small pure-Python package. sentence-transformers downloads all-MiniLM-L6-v2, about 90 MB, on first run. Outputs verified with rank_bm25 0.2.2 and sentence-transformers 5.4.1 on CPU.

A 15-product catalogue, one deliberately worded query

keyword_vs_semantic.py
import numpy as np
from rank_bm25 import BM25Okapi
from sentence_transformers import SentenceTransformer

products = [
    "Wireless earbuds with noise cancellation and 24 hour battery life",
    "Over-ear headphones with deep bass and a foldable design",
    "Bluetooth speaker, waterproof and portable, great for outdoor parties",
    "Laptop with 16GB RAM and fast SSD storage, built for programming",
    "Budget laptop for students, lightweight with long battery life",
    "Mechanical keyboard with RGB lighting and tactile switches",
    "Wireless mouse with an ergonomic shape that works on any surface",
    "Smartphone with a triple camera and all-day battery",
    "Budget smartphone with a big screen for watching videos",
    "Smartwatch that tracks your heart rate and sleep",
    "Fitness band with a step counter and water resistance",
    "External hard drive with 1TB of storage and fast USB-C transfer",
    "Portable power bank that charges your phone twice over",
    "Over-ear headphones built for long flights, with active noise cancelling",
    "4K webcam for video calls and live streaming",
]
query = "something to bring peace and quiet on a long flight"

# --- keyword search: BM25 ---
bm25 = BM25Okapi([p.lower().split() for p in products])
bm25_scores = bm25.get_scores(query.lower().split())
print("BM25 (keyword) top 3:")
for i in np.argsort(bm25_scores)[::-1][:3]:
    print(f"  {bm25_scores[i]:.2f}  {products[i]}")

# --- semantic search: embeddings + cosine similarity ---
model = SentenceTransformer("all-MiniLM-L6-v2")
doc_vecs = model.encode(products, normalize_embeddings=True)
q_vec = model.encode([query], normalize_embeddings=True)[0]
sem_scores = doc_vecs @ q_vec
print("\nSemantic (embeddings) top 3:")
for i in np.argsort(sem_scores)[::-1][:3]:
    print(f"  {sem_scores[i]:.2f}  {products[i]}")
Output
BM25 (keyword) top 3:
  2.09  Wireless mouse with an ergonomic shape that works on any surface
  1.71  Budget laptop for students, lightweight with long battery life
  1.63  Over-ear headphones built for long flights, with active noise cancelling

Semantic (embeddings) top 3:
  0.42  Over-ear headphones built for long flights, with active noise cancelling
  0.24  Bluetooth speaker, waterproof and portable, great for outdoor parties
  0.19  Budget laptop for students, lightweight with long battery life

The genuinely useful product — flight headphones with noise cancelling — ranks third under BM25 and first under semantic search. BM25's top pick, a wireless mouse, only got there because "a" and a couple of other low-value words happened to overlap with the query.

The walkthrough

BM25 scores by exact term overlap, weighted by rarity. The query shares almost no distinctive words with the correct product — "peace," "quiet" and "flight" versus "flights," "noise" and "cancelling." That near-miss on spelling is enough to sink BM25's ranking of the right answer, even though a person reads the two as an obvious match.

The semantic score comes from one dot product. Because both vectors are already unit-length (normalize_embeddings=True), doc_vecs @ q_vec computes cosine similarity for every product against the query in a single matrix multiply — no loop needed.

Why the semantic score for the winner is only 0.42, not close to 1.0. Cosine similarity between two different sentences, even a strong match, rarely approaches 1.0 in practice — that would require near-identical wording. Judge semantic scores by their ranking relative to each other, not against an absolute scale.

Common mistakes

Assuming semantic search is "better" and switching entirely. Keyword search still wins outright on exact codes, model numbers, names, and rare technical terms an embedding model may never have seen distinctly during training. A search for an exact product SKU is a case keyword search handles without effort, and embeddings can blur.

Comparing BM25 scores and cosine similarities on the same numeric scale. They are unrelated metrics with unrelated ranges — BM25 scores are unbounded and corpus-dependent; cosine similarity sits within [-1, 1]. Never mix them without a proper fusion method, covered in Hybrid search.

Not lower-casing and tokenising consistently for BM25. rank_bm25 does no text processing for you — feed it inconsistent casing or punctuation and the exact-match assumption it relies on breaks silently.

Try it yourself

Change the query to "RGB mechanical keyboard" — a query that shares exact words with a product — and re-run both searches. Notice both methods agree this time. Keyword and semantic search disagree most sharply on paraphrased, indirect queries, and agree readily on queries that already use the document's own words.

What to learn next

Researcher — Mathematics and papers.

BM25, formally

For query Q containing terms q_1, ..., q_n and document D:

text
BM25(D, Q) = sum over i=1..n of  IDF(q_i) * f(q_i, D) * (k1 + 1)
                                  ---------------------------------
                                  f(q_i, D) + k1 * (1 - b + b * |D| / avgdl)

Where:

  • f(q_i, D) — the number of times term q_i appears in document D.
  • |D| — the length of D in tokens; avgdl — the average document length in the corpus.
  • k1 — controls term-frequency saturation, typically 1.2–2.0 (extra occurrences of a word matter less and less).
  • b — controls length normalisation, typically 0.75 (longer documents are penalised for naturally containing more word matches).
  • IDF(q_i) = log( (N - n(q_i) + 0.5) / (n(q_i) + 0.5) + 1 ), where N is the corpus size and n(q_i) is how many documents contain q_i.

BM25 is a sparse, exact-match retrieval function: every dimension of its scoring corresponds to a literal vocabulary term, and a term contributes exactly zero unless it appears verbatim (after tokenisation) in both query and document.

Semantic search's failure modes, precisely

Dense embedding retrieval trades BM25's brittleness on synonyms for a different brittleness: out-of-distribution vocabulary. A model trained predominantly on general web and Wikipedia-style text (as most general-purpose sentence embedding models are) encodes rare technical jargon, product codes, and proper nouns less reliably, because it saw fewer or no comparable examples during training. This is well documented for exact identifiers — model numbers, legal citations, chemical compound names — where BM25's literal matching remains structurally more reliable than a learned continuous representation.

Why the two are usually combined, not chosen between

Robertson & Zaragoza's (2009) survey of the probabilistic relevance framework frames BM25 as optimising for relevance given observed term statistics — a well-calibrated, corpus-specific signal. Dense retrieval (Karpukhin et al., 2020, Dense Passage Retrieval, and the broader line of work it summarises) optimises for learned semantic proximity, calibrated by whatever data trained the encoder. Neither dominates the other across query types in the retrieval literature — short, jargon-heavy, or code-like queries tend to favour BM25; long, natural-language, paraphrased queries tend to favour dense retrieval. This is the empirical basis for hybrid search: combine both signals rather than commit to one.

Key references

  • Robertson, S. & Zaragoza, H. (2009). The Probabilistic Relevance Framework: BM25 and Beyond. Foundations and Trends in Information Retrieval 3(4).
  • Karpukhin, V. et al. (2020). Dense Passage Retrieval for Open-Domain Question Answering. arXiv:2004.04906
  • Reimers, N. & Gurevych, I. (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. arXiv:1908.10084

Current state

Production search systems increasingly run both retrieval methods in parallel and fuse their results rather than picking one, precisely because their failure modes are largely disjoint — the following lessons in this section build that pipeline piece by piece, from comparing individual encoders through combining full ranked lists.

What to learn next