How it works
How RAG works
RAG searches your documents first, then hands the best passages to the model so it answers from what it found instead of from memory.
- 4 min read
- Updated
On this page 9
An LLM on its own answers from memory, like a closed-book exam. RAG — retrieval-augmented generation — turns it into an open-book exam. The system looks up the right pages first, then answers using what it found. Here is the full pipeline, from raw documents to a cited answer.
The pipeline at a glance
BEFORE ANY QUESTION (ingestion, done once)
documents -> [1. chunk] -> [2. embed] -> [3. store in vector database]
WHEN A QUESTION ARRIVES
"What is our leave policy?"
|
v
[4. embed the question]
|
v
[5. search: find the nearest chunks]
|
v
[6. build a prompt: question + chunks]
|
v
[7. LLM writes the answer from the chunks]Stage 1 — chunk the documents
Whole documents are too long to search precisely, so each one is cut into chunks — passages of a few hundred tokens each. Chunk size is a real trade-off. Too small, and a chunk loses its context. Too large, and one chunk carries three topics, which blurs the search.
Stage 2 — turn each chunk into an embedding
Each chunk goes through an embedding model and comes out as an embedding: a long list of numbers placing the chunk in a "meaning space". Chunks about similar things land near each other. A chunk about annual leave lands near one about sick leave, far from one about GPU pricing.
Think of it as seating guests at a wedding. Similar people end up at the same table, so finding one relative means the rest are nearby.
Stage 3 — store the embeddings
The embeddings go into a vector database — a store built to answer one question fast: "which stored vectors are closest to this new one?" Each vector keeps a pointer back to its original text. This ingestion work happens once, ahead of time, and again whenever documents change.
Stage 4 — embed the question
Now a user asks something. The question goes through the same embedding model as the chunks did. Same model matters: both must land in the same meaning space, or nothing lines up.
Stage 5 — retrieve the nearest chunks
The vector database returns the handful of chunks closest to the question — the top 3 to 10, typically. This step is called retrieval, and it is where RAG succeeds or fails. If the right passage is not retrieved, the model never sees it, and no amount of clever prompting can recover.
Many production systems add a second pass here: a re-ranker, a model that reads the question against each candidate chunk and reorders them by actual relevance.
Stage 6 — build the prompt
The system assembles one prompt: instructions, the retrieved chunks pasted in as context, and the user's question. Something like: "Answer using only the passages below. Say so if they do not contain the answer."
This all has to fit in the model's context window — the fixed limit on how much text it reads at once. That limit is the reason for chunking in the first place.
Stage 7 — generate the answer
The LLM writes the answer, grounded in the pasted passages rather than its training memory. Because the system knows which chunks it used, it can cite sources — a link back to the exact document and section.
Grounding reduces hallucination — the model inventing plausible-sounding facts — but does not eliminate it. The model can still misread a passage, or blend two of them wrongly. Good RAG systems test for this with sets of questions whose correct answers are known.
Which lessons teach each stage
- The whole idea in depth: What is RAG?
- Stages 2 and 4, meaning as numbers: Embeddings
- Stages 3 and 5, the search engine underneath: Vector databases
- Stage 7, the model doing the writing: How LLMs actually work