Generative AI

What is RAG?

RAG means the model searches your documents first and then answers using what it found, instead of answering from memory.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works, in one picture
  5. The part people get wrong
  6. Where you have already seen it
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

RAG means the model looks things up before it answers.

RAG stands for retrieval-augmented generation — retrieval meaning search, augmented meaning helped along, generation meaning writing the answer.

The analogy you have already lived

Think of an exam where you were allowed to carry the textbook in. You did not have to memorise every date and formula. You found the right page, read it, and then wrote your answer in your own words.

A closed-book exam rewards memory. An open-book exam rewards knowing where to look.

An LLM on its own is sitting a closed-book exam. RAG hands it the textbook.

Why it exists

A model learns from text that was collected months or years before you use it. Four problems follow directly from that.

It goes stale. Anything that happened after training ended is invisible to it.

It never saw your documents. Your company's leave policy, your product manual and your class notes were not on the public internet.

It invents things. When a model does not know, it produces the most plausible-sounding text instead of stopping. That is hallucination.

You cannot check it. A plain answer comes with no source, so you have no way to verify it.

RAG attacks all four with one move. Find the right passage first. Paste it into the question. Ask the model to answer using only that passage, and to cite it.

How it works, in one picture

   your question
        |
        v
  [ search your documents ]  ->  the three closest passages
        |
        v
  [ build a new prompt ]
        "Answer using ONLY this text:
           <passage 1>
           <passage 2>
           <passage 3>
         Question: ..."
        |
        v
     [ model ]
        |
        v
   an answer, plus the passages it came from

The model has not learned anything new. It has been handed the right page at the right moment.

The part people get wrong

RAG does not change the model at all. Nothing is trained, nothing is stored inside it, no weights move.

Every conversation starts fresh, and every conversation gets a fresh search. If you replace a document today, the answers change today. That is a feature, and it is why RAG beats retraining for anything that updates often.

Where you have already seen it

  • A bank's help chat answering from its own policy pages.
  • Search engines that show an AI summary with links underneath.
  • "Chat with your PDF" tools that answer questions about a file you uploaded.
  • A support bot that quotes your actual order status.

The honest part

RAG is not a solved problem, and the failure mode is quiet.

If the search step brings back the wrong passage, the model will answer confidently from the wrong passage. The answer looks well-sourced and is wrong. Most RAG systems that disappoint in production are failing at search, not at writing.

So the work is mostly unglamorous. You cut documents into sensible pieces. You check that the right piece comes back. Then you measure it. Anyone who calls RAG a weekend feature has not run one for six months.

Remember this

  • RAG searches first, then answers using what it found.
  • The model is unchanged. Only the prompt gets richer.
  • When RAG fails, blame the search step before blaming the model.

What to learn next

Developer — Code and libraries.

Setup

bash
# the core example uses the standard library only
python3 --version

RAG has three moving parts: index, retrieve, generate. Two of them are ordinary information retrieval and you can build them without any model at all. Doing that first is the fastest way to understand where RAG actually breaks.

A complete retriever in forty lines

tiny_rag.py
import re
from collections import Counter

DOCS = {
    "refunds":  "A refund is sent back to your UPI app within five working days.",
    "delivery": "Delivery inside India takes three to five working days.",
    "returns":  "You may return an unused item within thirty days of delivery.",
    "warranty": "Every laptop has a one year warranty covering manufacturing defects.",
}

STOP = {"a", "an", "the", "is", "are", "to", "of", "in", "and", "you", "your",
        "my", "i", "how", "does", "do", "will", "get", "for", "it", "within",
        "many", "when"}

def words(text):
    return [w for w in re.findall(r"[a-z]+", text.lower()) if w not in STOP]

INDEX = {name: Counter(words(body)) for name, body in DOCS.items()}

def retrieve(question, k=2):
    q = Counter(words(question))
    # score = how many query words also appear in the document
    scored = [(sum(min(n, INDEX[name][w]) for w, n in q.items()), name) for name in INDEX]
    scored.sort(key=lambda pair: (-pair[0], pair[1]))   # name breaks ties, so runs repeat
    return scored[:k]

def build_prompt(question, hits):
    context = "\n".join(f"[{name}] {DOCS[name]}" for _, name in hits)
    return ("Answer using ONLY the context. If it is not there, say you do not know.\n\n"
            f"CONTEXT:\n{context}\n\nQUESTION: {question}\nANSWER:")

q1 = "How many working days for delivery inside India?"
hits = retrieve(q1)
print("top matches for a well-worded question:")
for score, name in hits:
    print(f"{score}  {name}")
print()
print(build_prompt(q1, hits))
print()
print("different wording, same meaning:", retrieve("Will I be reimbursed for a late order?"))
Output
top matches for a well-worded question:
5  delivery
2  refunds

Answer using ONLY the context. If it is not there, say you do not know.

CONTEXT:
[delivery] Delivery inside India takes three to five working days.
[refunds] A refund is sent back to your UPI app within five working days.

QUESTION: How many working days for delivery inside India?
ANSWER:

different wording, same meaning: [(0, 'delivery'), (0, 'refunds')]

Read that output carefully — it contains both lessons

The first query worked. delivery scored 5 and came back first. The prompt that got built is a real RAG prompt: an instruction, a fenced context with source labels, and the question last.

Notice that refunds also came back, scoring 2 on the words "working" and "days". It is irrelevant to the question. Retrieval always brings noise, and the instruction "answer using ONLY the context" has to survive it. Retrieving more passages is not free.

The second query failed completely. "Will I be reimbursed for a late order?" is asking about refunds. Every document scored zero, and the function still returned two documents, because it was told to return two.

That is the single most important thing on this page. The word "reimbursed" never appears in the refunds document, which says "refund". Word matching has no idea those are related. Neither does it know that "late" relates to delivery.

This is exactly the gap embeddings close. An embedding puts "reimbursed" and "refund" near each other in number space, so the search finds one from the other. See vector databases for the mechanics.

Line-by-line, for the parts that trip people up

Counter[missing_key] returns 0 without inserting the key. That keeps INDEX from silently growing every time you score a query, which is a real bug in hand-rolled versions of this.

sum(min(n, INDEX[name][w]) for w, n in q.items()) clamps the contribution of a repeated query word to how often it actually appears in the document. Without the min, typing "days days days days" would rank a document higher.

scored.sort(key=lambda pair: (-pair[0], pair[1])) sorts by score descending, then by name ascending. The second key is not decoration. Without it, two documents on the same score would come back in dictionary order, and your test suite would fail whenever you added a document.

retrieve returning k results even when every score is zero is a deliberate demonstration of a real production bug. Fix it in your own code with a floor:

python
hits = [(s, n) for s, n in retrieve(question) if s > 0]
if not hits:
    return "I could not find anything about that in the documents."

Refusing to answer is a valid answer. Systems without that branch hallucinate for a living.

The generation step

Once the prompt is built, generation is an ordinary model call. Using a local Ollama server so nothing leaves your machine:

python
import json, urllib.request

payload = {"model": "llama3.2", "prompt": build_prompt(q1, hits), "stream": False}
req = urllib.request.Request(
    "http://localhost:11434/api/generate",
    data=json.dumps(payload).encode(),
    headers={"Content-Type": "application/json"},
)
with urllib.request.urlopen(req) as r:
    print(json.loads(r.read())["response"])

No output block here, deliberately. The answer text varies by model, by version and between runs. An invented sample would teach you to expect something that will not happen.

Chunking, the part nobody warns you about

Real documents are longer than one sentence, so they get cut into chunks — pieces small enough to fit several into a prompt.

Getting this wrong ruins everything downstream, and there is no universal right answer. Some practical starting points:

  • Around 300 to 800 tokens per chunk for prose. Smaller for dense reference material.
  • Overlap neighbouring chunks by 10 to 20 percent, so a fact split across a boundary survives in at least one chunk.
  • Split on structure first — headings, sections, list items — and only fall back to fixed sizes.
  • Keep the heading path inside the chunk text. A chunk that reads "must be renewed annually" is useless without "Section 4.2, Trade licences" attached to it.

Common mistakes

Measuring the wrong thing. Teams evaluate the final answer and conclude the model is weak. Measure retrieval separately: for a set of questions with known correct chunks, what fraction of the time is the right chunk in the top k? If that number is 0.6, no prompt will save you.

Retrieving too much. More context is not better. Accuracy on facts placed in the middle of a long context measurably drops. Three good chunks beat twenty mediocre ones.

No citations. Ask the model to tag each claim with the source label you put in the context. It makes wrong answers detectable by a human in seconds.

Rebuilding the whole index on every change. Store a hash per source document and reindex only what moved. This becomes painful at exactly the moment your corpus gets interesting.

Assuming a long context window removes the need for RAG. Pasting a whole handbook into every request costs money on every request and gets slower as it grows. Retrieval is a cost-control mechanism as much as an accuracy one.

Try it yourself

Add a fifth document about cancelling an order. Then write five questions a real customer would type, in the words a real customer would use. Count how many return the right document. That number is your retrieval recall, and it is the honest health metric for any RAG system you build.

What to learn next

  • Vector databases — replacing word matching with meaning matching.
  • Embeddings — the representation that makes it possible.
  • LlamaIndex — a framework that wires these steps together.

Researcher — Mathematics and papers.

The original formulation

Lewis et al. (2020) defined RAG as marginalising the generator over retrieved documents:

p(y | x)  ≈  ∑  p_η(z | x) · p_θ(y | x, z)
            z ∈ top-k(x)

x is the query, y the output sequence, z a retrieved passage, η the retriever parameters and θ the generator parameters. Two variants were proposed. RAG-Sequence conditions the whole output on one z. RAG-Token marginalises per generated token, allowing different passages to support different parts of one answer.

Both were trained end to end, backpropagating into the query encoder while the document index stayed fixed. Almost no production system does this today. The industry standard is a frozen retriever and a frozen generator joined by prompt concatenation, which is strictly weaker than the paper and vastly easier to operate.

Fusion-in-Decoder (Izacard and Grave, 2020) is the important middle ground: encode each passage independently, concatenate in the decoder. Cost grows linearly in k rather than quadratically, which is why it scales to large k.

Retrieval, three families

Sparse lexical. BM25 remains a genuinely strong baseline:

score(q,d) = ∑  IDF(t) · ( f(t,d) · (k1 + 1) )
             t∈q          ---------------------------------------
                          f(t,d) + k1 · (1 − b + b · |d| / avgdl)

IDF(t) = ln( 1 + (N − n_t + 0.5) / (n_t + 0.5) )

f(t,d) is the frequency of term t in document d. |d| is document length and avgdl the mean length over the corpus. N is the number of documents and n_t the number containing t. k1 controls term-frequency saturation, usually 1.2 to 2.0. b controls length normalisation, usually 0.75.

Dense bi-encoder. DPR (Karpukhin et al., 2020) trains separate query and passage encoders with an inner-product objective and in-batch negatives. Passages are embedded once, offline. Query cost is one forward pass plus an approximate nearest-neighbour lookup.

Late interaction. ColBERT (Khattab and Zaharia, 2020) keeps per-token embeddings and scores by summing per-query-token maxima over document tokens. Higher quality than a single-vector bi-encoder, at a substantially larger index.

Hybrid retrieval

Sparse and dense fail on different queries. Sparse handles exact identifiers, product codes and rare proper nouns. Dense handles paraphrase and synonymy. Reciprocal rank fusion (Cormack et al., 2009) combines ranked lists without needing calibrated scores:

RRF(d)  =  ∑  1 / ( k + r_i(d) )
           i

r_i(d) is the rank of document d in list i, and k is a smoothing constant conventionally set to 60. The absence of score normalisation is the point: raw BM25 scores and cosine similarities are not on comparable scales, and fusing ranks sidesteps that entirely.

Reranking

A cross-encoder scores the query and passage jointly in a single forward pass, which is far more accurate than a bi-encoder and far too slow to run over a corpus. The standard cascade retrieves k ≈ 100 cheaply, then reranks to the top 5 with O(k) cross-encoder passes. This is usually the highest quality-per-engineering-hour change available to a struggling RAG system.

Position effects

Liu et al. (2023) documented a U-shaped accuracy curve over the position of relevant information within a long context: strong at the beginning, strong at the end, measurably weaker in the middle. The practical consequences are concrete. Order retrieved passages so the most relevant sit at the edges, and treat "increase k" as a change that can reduce accuracy.

Evaluation

Separate the stages or you will chase the wrong bug.

  • Retrieval: recall@k, MRR, nDCG@10 against labelled query-passage pairs.
  • Grounding: is every claim in the answer entailed by the retrieved context? Natural-language-inference models or an LLM judge with a strict rubric both work; both need a human-labelled calibration set.
  • Answer quality: exact match or F1 for extractive questions, pairwise preference for open-ended ones.

RAGAS (Es et al., 2023) packages faithfulness, answer relevance and context relevance as reference-free metrics. Useful for regression detection, weak as ground truth. Do not ship on it alone.

Active directions

Self-RAG (Asai et al., 2023) trains the model to emit reflection tokens deciding when to retrieve and whether the retrieved evidence supports its own draft. HyDE (Gao et al., 2022) embeds a hypothetical answer generated from the query rather than the query itself, closing the vocabulary gap between short questions and long passages at the cost of one extra generation.

The open problem remains upstream of all of it. Retrieval failure dominates the error budget in almost every deployed system, and no amount of generator quality compensates for a missing passage.

Papers

What to learn next

  • Vector databases — approximate nearest neighbour search in practice.
  • Embeddings — bi-encoder training and representation quality.
  • Fine-tuning — when to adapt weights instead of retrieving.

What to learn next

These follow on from what you just read.

  • Generative AI

    Vector databases

    A vector database stores text as numbers that capture meaning, then finds the closest matches to your question in milliseconds.

  • Generative AI

    Fine-tuning

    Fine-tuning continues training an already-trained model on your own examples, which changes how it behaves — and is the wrong tool for most problems beginners reach for it with.

  • Generative AI

    LoRA

    LoRA fine-tunes a large model by freezing it and training a small add-on beside it, which cuts the memory cost enormously and lets you swap behaviours like plug-in packs.