How it works
How search ranking works
A search engine reads the web ahead of time into an index, understands your query, pulls thousands of candidate pages, and ranks them with hundreds of signals — all before you finish blinking.
- 4 min read
- Updated
On this page 7
You type three words and get ten links from billions of pages, in about half a second. The trick is that almost none of the work happens when you search. The web was read, sorted and indexed long before, so your query is a lookup, not a hunt. Here is the pipeline on both sides of the search box.
The pipeline at a glance
BEFORE YOU SEARCH (continuous, in the background)
the web -> [1. crawl] -> [2. index] "which pages contain which words"
WHEN YOU SEARCH
"best phone under 20000"
|
v
[3. understand the query] spelling, synonyms, intent
|
v
[4. retrieve candidates] thousands of maybe-relevant pages
|
v
[5. rank] hundreds of signals -> ordered list
|
v
ten blue links (and what you click teaches stage 5)Stage 1 — crawl: read the web ahead of time
Programs called crawlers fetch pages, follow the links on them, fetch those pages, and never stop. Busy sites get revisited hourly, sleepy ones monthly. Everything found is stored and parsed.
This is why a brand-new page cannot rank tonight: until a crawler has seen it, the search engine does not know it exists.
Stage 2 — index: build the reverse phone book
Searching billions of stored pages one by one is impossible. So the engine builds an inverted index: for every word, the list of pages containing it. A normal book maps page to words; the index at the back maps word to pages. Same idea, web-sized.
Now "which pages mention 'mango'" is a single lookup, not a scan. Alongside the text, the index stores each page's reputation signals — most famously how many other sites link to it, the idea behind PageRank.
Stage 3 — understand the query
Your query gets cleaned and interpreted before anything is fetched. Spelling is corrected. Synonyms are expanded — "phone" should also match "smartphone" and "mobile". Intent is classified: "paneer recipe" wants pages, "weather" wants a widget, "pizza near me" wants maps.
Modern engines also convert the query into an embedding — a list of numbers capturing its meaning — so pages can match by sense, not spelling. "How to fix a flat tyre" can retrieve a page titled "repairing a puncture" with zero words in common.
Stage 4 — retrieve candidates
Now the funnel starts. Fast, cheap lookups — keyword matches from the inverted index, plus nearest-neighbour matches in embedding space — pull a few thousand plausible pages from the billions. Precision does not matter yet. The only sin at this stage is missing the right page entirely, because whatever is not retrieved can never be ranked.
Stage 5 — rank the candidates
The expensive machinery runs only on those few thousand. A ranking model scores each page against your query using hundreds of signals. Does the title match? Is the content substantial? Do trusted sites link to it? Is it fresh, and does it load well on a phone? Learned models blend the signals — the field is literally called learning-to-rank — and the sorted result becomes your page of links.
Then the loop quietly closes. Clicks, skips and quick bounces back to the results page become training data for tomorrow's ranker.
The honest caveat: ranking is adversarial. An entire industry — SEO — exists to reverse-engineer these signals, which is why engines keep the exact recipe secret and change it constantly. Every ranking is one move in a long game between the engine and people who want to be ranked first.
Which lessons teach each stage
- Stage 3, meaning as numbers: Embeddings
- Stage 4, fast nearest-neighbour retrieval: Vector databases
- Stage 5, judging whether a ranker improved: Model evaluation
- The same retrieve-then-rank funnel inside AI answers: What is RAG?