AI glossary

Reranking

In one sentence Reranking is a second, more careful scoring pass over retrieved search results, promoting the truly relevant chunks before they reach the model.

By Updated

Reranking takes the rough top results from a fast search and re-scores them with a slower, smarter model, so the best few rise to the top.

It is exactly how hiring works. A quick CV screen cuts 500 applicants to 30 — fast, keyword-ish, and it wrongly ranks some strong candidates twentieth. Then interviews examine those 30 properly and produce the real top 5. Nobody interviews all 500 (too slow), and nobody hires straight from the CV screen (too crude). Two stages, each doing what it is good at.

In a RAG pipeline: vector-database search fetches, say, the top 50 chunks by embedding similarity — fast because chunks were embedded ahead of time, coarse because each chunk was compressed to one vector before knowing the question. The reranker (a cross-encoder: a model that reads query and chunk together, token by token) then scores those 50 pairs properly and reorders. The model receives the reranked top 5-10.

question → vector search → 50 rough candidates
         → reranker reads (question + chunk) pairs → true top 5 → prompt

The gain is real and cheap to test: rerankers routinely lift retrieval accuracy noticeably, and since retrieval quality caps the whole system, that lift passes straight through to answers. Off-the-shelf options: Cohere Rerank, Jina, BGE-reranker, plus open cross-encoders in sentence-transformers. Costs: extra latency per query (mind the candidate count), and one more component to monitor.

Where to go next