Semantic Search and Reranking

Query rewriting and expansion

Add extra terms to a query using what its own top results already have in common. It genuinely improves recall, and it can also backfire by pulling in an off-topic result — this lesson shows both happening in the same example.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Query expansion adds extra words to a search, based on what the first results already have in common. It catches things the original wording missed.

Ask a shopkeeper for "something for a headache." A good one does not hand you only the one item literally labelled that way. They also point you toward the balm and the paracetamol sitting nearby. Those are related to what you actually need, even though you never named them.

Query expansion teaches a search system to do the same thing automatically. After seeing what your first search turns up, it notices related words in those results. It quietly adds them to your query before searching again.

Why it exists

A short, vague query — "office gear" say — gives a search system very little to work with. Somewhere in the collection sits a document, worded quite differently, that would answer the same need perfectly well. The original query's exact words never reach it. Expansion tries to bridge that gap, using the search system's own first attempt as a clue.

How it works

query: "office gear"                    (vague, only two words)
        |
        v
run the search once, look at the TOP FEW results
  -> they mention: "keyboard", "waterproof", "bluetooth", "tactile"
        |
        v
add the most distinctive of those words back into the query
  -> "office gear keyboard tactile ..."
        |
        v
search AGAIN with the expanded query
  -> catches documents the original two words alone would have missed

This version is called pseudo-relevance feedback. The "extra clues" come from the search's own results, not a human. It is "pseudo" because nobody confirmed those first results were relevant before using them as feedback.

A real example you have seen

Search "cheap flights" on a travel site. It also quietly matches listings that say "budget airfare." No exact word needs to match. Search systems commonly expand a query with related terms. Those terms are pulled from similar past searches, or from what the top results tend to contain.

Remember this

  • Query expansion adds extra words to a search based on its own initial results, to catch documents the original wording missed.
  • Pseudo-relevance feedback assumes the top few results are relevant, without checking — that assumption is not always true.
  • Expansion can genuinely help recall, and it can also pull the search off-topic. Both are demonstrated in this lesson's code.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install sentence-transformers scikit-learn

Outputs verified with sentence-transformers 5.4.1 and scikit-learn 1.7.2 on CPU.

A vague query, expanded from its own top-3 results

query_expansion.py
import numpy as np
from sentence_transformers import SentenceTransformer
from sklearn.feature_extraction.text import TfidfVectorizer

products = [
    "Wireless earbuds with noise cancellation and 24 hour battery life",
    "Over-ear headphones with deep bass and a foldable design",
    "Bluetooth speaker, waterproof and portable, great for outdoor parties",
    "Laptop with 16GB RAM and fast SSD storage, built for programming",
    "Budget laptop for students, lightweight with long battery life",
    "Mechanical keyboard with RGB lighting and tactile switches",
    "Wireless mouse with an ergonomic shape that works on any surface",
    "Smartphone with a triple camera and all-day battery",
    "Budget smartphone with a big screen for watching videos",
    "Smartwatch that tracks your heart rate and sleep",
    "Fitness band with a step counter and water resistance",
    "External hard drive with 1TB of storage and fast USB-C transfer",
    "Portable power bank that charges your phone twice over",
    "Over-ear headphones built for long flights, with active noise cancelling",
    "4K webcam for video calls and live streaming",
]
model = SentenceTransformer("all-MiniLM-L6-v2")
doc_vecs = model.encode(products, normalize_embeddings=True)


def search(q):
    q_vec = model.encode([q], normalize_embeddings=True)[0]
    scores = doc_vecs @ q_vec
    return np.argsort(scores)[::-1], scores


query = "office gear"
order, scores = search(query)
print("original query:", repr(query))
for rank, i in enumerate(order[:5], start=1):
    print(f"  {rank}. ({scores[i]:.3f}) {products[i]}")
mouse_rank = int(np.where(order == 6)[0][0]) + 1
print(f"  -> wireless mouse is currently rank {mouse_rank}")

# pseudo-relevance feedback: assume the top 3 are relevant, pull their standout words into the query
top3_text = " ".join(products[i] for i in order[:3])
tfidf = TfidfVectorizer(stop_words="english")
tfidf_scores = tfidf.fit_transform([top3_text]).toarray()[0]
words = tfidf.get_feature_names_out()
feedback_terms = [words[i] for i in tfidf_scores.argsort()[::-1][:5]]
expanded_query = query + " " + " ".join(feedback_terms)

print("\nfeedback terms pulled from the top 3 results:", feedback_terms)
print("expanded query:", repr(expanded_query))

order2, scores2 = search(expanded_query)
for rank, i in enumerate(order2[:5], start=1):
    print(f"  {rank}. ({scores2[i]:.3f}) {products[i]}")
mouse_rank2 = int(np.where(order2 == 6)[0][0]) + 1
print(f"  -> wireless mouse is now rank {mouse_rank2}")
Output
original query: 'office gear'
  1. (0.236) Mechanical keyboard with RGB lighting and tactile switches
  2. (0.191) Budget laptop for students, lightweight with long battery life
  3. (0.167) Bluetooth speaker, waterproof and portable, great for outdoor parties
  4. (0.160) Budget smartphone with a big screen for watching videos
  5. (0.156) Fitness band with a step counter and water resistance
  -> wireless mouse is currently rank 9

feedback terms pulled from the top 3 results: ['waterproof', 'tactile', 'bluetooth', 'budget', 'great']
expanded query: 'office gear waterproof tactile bluetooth budget great'
  1. (0.530) Bluetooth speaker, waterproof and portable, great for outdoor parties
  2. (0.422) Mechanical keyboard with RGB lighting and tactile switches
  3. (0.367) Fitness band with a step counter and water resistance
  4. (0.308) Wireless mouse with an ergonomic shape that works on any surface
  5. (0.292) Wireless earbuds with noise cancellation and 24 hour battery life
  -> wireless mouse is now rank 4

The wireless mouse, a genuinely relevant "office gear" item, climbed from 9th to 4th — expansion worked, exactly as intended. Look at rank 1, though: the bluetooth speaker, which is not office gear at all, climbed from 3rd to 1st. It only got into the "assumed relevant" top-3 in the first place because the original vague query was weak — and expansion then trusted that mistake and amplified it.

The walkthrough

This is genuinely one example showing both outcomes at once. The mouse's improvement is expansion working as designed: real office-related words ("tactile," from the keyboard result) pulled a relevant document up. The speaker's rise is query drift — expansion trusting an irrelevant result that slipped into the "assumed relevant" top-3, and reinforcing it instead of correcting it.

top3_text is the entire feedback signal, and it is only as good as the original weak query's top-3. Pseudo-relevance feedback has no way to check whether those top-3 documents are actually relevant before using them — it takes the search engine's own imperfect first guess at face value.

TfidfVectorizer on a single merged string finds the standout words of that top-3 set. Only one "document" (the merged top-3 text) is passed in, so every surviving word after stop-word removal gets scored, and the highest-weighted ones become the feedback terms.

Common mistakes

Trusting expansion blindly on a query that already worked well. If the original top result was already the right one, expansion has nothing useful to add and only risks introducing drift, as it does with the speaker here.

Pulling feedback terms from too wide a slice of results. The wider the "assumed relevant" set, the more likely it includes something off-topic like the speaker above. Top-3 to top-5 is a common practical range; going much wider raises drift risk substantially.

Not weighting the original query terms more heavily than the added feedback terms. In this toy example the expanded query treats every word equally. Rocchio's classic feedback algorithm and most production systems instead down-weight the added terms relative to the original query, precisely to limit how far a single bad feedback term can drag the result.

Try it yourself

Try the query "gift for someone who exercises" instead, where the original top result is already correct (a fitness band). Confirm expansion is unnecessary there, and notice whether it makes the ranking better, worse, or barely different — a useful check on when expansion is worth the risk.

What to learn next

Researcher — Mathematics and papers.

Rocchio's algorithm, the classical formulation

Rocchio (1971) formalised relevance feedback for the vector-space model as a linear combination of the original query vector with the centroids of known relevant and non-relevant documents:

text
q_new = alpha*q_0 + beta * (1/|D_r|) * sum over d in D_r of d
                   - gamma * (1/|D_nr|) * sum over d in D_nr of d

Where q_0 is the original query vector, D_r and D_nr are the sets of known relevant and non-relevant documents, and alpha, beta, gamma are weights (commonly alpha=1, beta=0.75, gamma=0.15 in classic settings) controlling how strongly feedback shifts the query. True relevance feedback requires human judgments to populate D_r and D_nr.

Pseudo-relevance feedback and the drift problem

Pseudo-relevance feedback (PRF), demonstrated in the developer block, substitutes the top-k retrieved documents for D_r without human confirmation, and typically omits an explicit D_nr term entirely (Xu & Croft, 1996, Query Expansion Using Local and Global Document Analysis, is a foundational treatment). Because this assumption is unverified, PRF inherits whatever errors the initial retrieval made. The failure mode demonstrated above — an irrelevant top-3 document contributing terms that then dominate the expanded query — is precisely what the information retrieval literature refers to as query drift: expansion terms pulling the query semantically away from the user's actual intent, most likely exactly when the initial retrieval was already weak and needed less trust placed in it, not more.

Mitra, Singhal & Buckley (1998) studied this trade-off directly, finding PRF's average benefit across many queries was real but came with meaningfully higher variance — some queries improved substantially, a minority got measurably worse, consistent with the single example shown here generalising into an aggregate pattern.

Modern approaches

RM3 / relevance models (Lavrenko & Croft, 2001) formalise PRF probabilistically, estimating a language model over the query's relevant-document distribution from the top-k results and interpolating it with the original query's language model, with an explicit interpolation weight controlling how much to trust the feedback — a direct, tunable answer to the drift problem.

Neural and LLM-based query rewriting has largely supplemented classical PRF in modern systems: a language model, prompted with the original query (and sometimes the initial results), generates an expanded or reformulated query directly, conditioned on much broader world knowledge than a single corpus's top-3 results can provide. This does not eliminate the drift risk in principle — a language model can also confidently add the wrong context — but it draws on a far richer source of candidate expansion terms than pseudo-relevance feedback's narrow, self-referential loop.

Key references

  • Rocchio, J. (1971). Relevance Feedback in Information Retrieval. In The SMART Retrieval System.
  • Xu, J. & Croft, W.B. (1996). Query Expansion Using Local and Global Document Analysis. SIGIR.
  • Lavrenko, V. & Croft, W.B. (2001). Relevance-Based Language Models. SIGIR.
  • Mitra, M., Singhal, A. & Buckley, C. (1998). Improving Automatic Query Expansion. SIGIR.

Current state

Pseudo-relevance feedback remains a standard, low-cost technique for improving average recall in classical and dense retrieval systems, with well-understood drift risk that production systems typically mitigate by keeping the feedback pool narrow, down-weighting added terms relative to the original query, and monitoring aggregate metrics (not single-query outcomes) to confirm net benefit — exactly the caution this lesson's worked example is meant to make concrete.

What to learn next