Generative AI

Hallucination

A hallucination is a confident false statement, and it follows directly from how a language model predicts text — it is not a bug that someone forgot to fix.

On this page 11
  1. Why this happens
  2. There is no list of facts inside it to check
  3. It was also taught that answering is good
  4. How it works, in one picture
  5. It is not a bug someone forgot to fix
  6. Where you have already seen it
  7. How to reduce it
  8. How to spot it
  9. The honest part
  10. Remember this
  11. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A hallucination is when a model says something false, in a confident, well-written sentence.

Think about asking a stranger for directions in a city you do not know. Some people say "sorry, no idea". Others point firmly down a road they are not sure about, because pointing feels more helpful than shrugging. You walk twenty minutes before you find out.

A language model is always the second kind of person. It does not know that it does not know.

The answer arrives in the same calm, fluent voice whether it is right or wrong. That is what makes it dangerous.

Why this happens

This section is the whole lesson. Read it slowly.

A language model does exactly one thing: it guesses what piece of text comes next. It does that over and over until an answer has appeared. That is the entire machine — see how LLMs work.

Now look at what is missing from that description. There is no step where it checks whether the sentence is true. No lookup, no database, no verification. Nowhere in the process does anything ask that question.

So a true sentence and a false-but-natural-sounding sentence look identical to the model. Both are ordinary text. Both flow well. It has no way to prefer one over the other. The thing it measures is "does this read like language", not "is this so".

There is no list of facts inside it to check

People imagine the model holds a giant table of facts and sometimes reads the wrong row. That picture is wrong, and dropping it explains a lot.

What the model learned is spread thinly across billions of numbers. Nothing is stored as a clean, separate item. There is no row for your city's pin code that could be looked up and confirmed.

Facts that appeared many times in training are represented strongly and come out reliably. Facts that appeared once or twice are faint. When the model reaches for a faint one, it produces the shape of an answer with plausible details filled in.

It was also taught that answering is good

During training, replies that answered the question were preferred by people. "I am not sure" is rarely the reply anyone marks as best.

So the model learns to produce an answer. Guessing scores better than admitting ignorance, in the same way a guessed answer scores better than a blank on a multiple-choice paper. Nobody intended that lesson. It falls out of how the scoring worked.

How it works, in one picture

   your question
        |
        v
   [ what text is likely to come next? ]  ──►  a piece of text
        |                                            |
        └──────────  add it and repeat  ◄────────────┘
        |
        v
   a fluent, complete, confident answer

   nowhere in this loop:  [ is this true? ]

It is not a bug someone forgot to fix

This matters, so it gets its own heading.

A bug is a mistake in the code that a careful engineer can find and repair. There is no such line here. Hallucination comes out of the design of the machine, the same way a shadow comes out of a lamp.

It can be reduced a great deal, and the rest of this page is about how. It cannot be removed while the model works by predicting likely text.

So treat any claim that some model "does not hallucinate" as a marketing sentence. Newer models hallucinate less on common questions. All of them still do it, and often the confident tone gets better faster than the accuracy does.

Where you have already seen it

  • Court cases. In 2023 a New York lawyer filed a brief containing case citations produced by a chatbot. The cases did not exist. He was sanctioned. The citations looked perfect: real-sounding names, plausible numbers, correct formatting.
  • Books that were never written. Ask for reading suggestions on a narrow topic. You often get a real author, a real publisher, and a title nobody ever wrote.
  • Code. It invents functions that sound exactly like something the library should have. The name is so plausible that developers file bug reports asking why it is missing.
  • Local details. Bus numbers, shop timings, office addresses, helpline numbers. These appear rarely in training text, so they come out faint and get filled in.

How to reduce it

These are ordered by how much they help.

1. Give it the source and tell it to stay inside. Paste the document, or use RAG to fetch it. Then say: answer using only the text above; if it is not there, say you do not know. This is by far the biggest single improvement available to you.

2. Give it permission to refuse. Add a line to your request: if you are not sure, say so instead of guessing. It will not always obey, but it obeys far more often than a model that was never offered the option.

3. Ask for the exact quote. Make it copy the sentence its answer came from. A model that has to quote is much harder to drift, and you can check the quote in seconds.

4. Turn the randomness down. For factual work, set temperature low. That does not make it truthful, but it stops it wandering into unlikely text. See temperature and sampling.

5. Break big questions into small ones. One question with five parts gives five chances to invent something, and one invented part will drag the rest along with it.

6. Never let it do arithmetic or dates in its head. Hand those to a calculator or to code. This is a solved problem and it does not need a model.

How to spot it

  • Check the crunchy bits first. Names, numbers, dates, links, citations, section references. Fluent text is where invention hides, and specific details are where it is easiest to catch.
  • Search for the exact claim. If a case, a paper or a product does not exist, one search will usually show it.
  • Ask the same question in a fresh chat. Two or three times. If the details change between runs, the model is guessing. Stable answers are not proof of truth, but changing answers are strong proof of guessing.
  • Ask where it came from. The source it names can be invented too. Ask anyway, because a fabricated source is often easier to check than the claim itself.
  • Be most suspicious where you know least. You catch errors in your own field instantly. The same error rate is there in every field you cannot check.

The honest part

Confidence tells you nothing. Neither does good writing, formatting, or the polite hedging language models produce.

There is no reliable "how sure are you" you can ask for. Asked to rate its own confidence, a model produces the text that a confident answer looks like. That is a different thing from being right.

And this is not solved. It is an active research problem with real progress and no ending in sight. Anyone who tells you otherwise is selling something.

Remember this

  • A model predicts likely text. There is no truth check anywhere inside it.
  • Hallucination is a consequence of the design, not a bug awaiting a patch.
  • The strongest defence is giving it the source and making it quote from it.

What to learn next

Developer — Code and libraries.

Setup

bash
python3 --version

The first three examples use the standard library only. Nothing is downloaded, nothing needs a key, and every number below is reproducible.

The mechanism, in twenty lines

Here is a tiny language model. It learns which word follows which word, from three true sentences. Then we ask it to score every sentence it can produce.

no_truth_check.py
from collections import defaultdict

# Three true sentences. This is our model's entire world.
FACTS = [
    "the capital of france is paris",
    "the capital of japan is tokyo",
    "the capital of kenya is nairobi",
]

# Count which word follows which word. This is the whole "model".
follows = defaultdict(lambda: defaultdict(int))
for line in FACTS:
    w = line.split()
    for a, b in zip(w, w[1:]):
        follows[a][b] += 1

def prob(word, nxt):
    row = follows[word]
    return row[nxt] / sum(row.values())

def sentence_prob(words):
    p = 1.0
    for a, b in zip(words, words[1:]):
        p *= prob(a, b)
    return p

print("what the model thinks can follow each word:")
for word in ["of", "is"]:
    print(f"  after '{word}': {dict(follows[word])}")

print()
print("every sentence this model can produce, and how likely it thinks each one is:")
for country in ["france", "japan", "kenya"]:
    for capital in ["paris", "tokyo", "nairobi"]:
        s = f"the capital of {country} is {capital}"
        true = (country, capital) in [("france", "paris"), ("japan", "tokyo"), ("kenya", "nairobi")]
        mark = "TRUE " if true else "FALSE"
        print(f"  {mark}  p = {sentence_prob(s.split()):.4f}   {s}")
Output
what the model thinks can follow each word:
  after 'of': {'france': 1, 'japan': 1, 'kenya': 1}
  after 'is': {'paris': 1, 'tokyo': 1, 'nairobi': 1}

every sentence this model can produce, and how likely it thinks each one is:
  TRUE   p = 0.1111   the capital of france is paris
  FALSE  p = 0.1111   the capital of france is tokyo
  FALSE  p = 0.1111   the capital of france is nairobi
  FALSE  p = 0.1111   the capital of japan is paris
  TRUE   p = 0.1111   the capital of japan is tokyo
  FALSE  p = 0.1111   the capital of japan is nairobi
  FALSE  p = 0.1111   the capital of kenya is paris
  FALSE  p = 0.1111   the capital of kenya is tokyo
  TRUE   p = 0.1111   the capital of kenya is nairobi

Read that output twice. It is the whole lesson in nine lines.

The model was trained only on true sentences. Every one of its nine possible outputs scores identically. Six of them are false.

Nothing went wrong here. No bug, no bad data, no missing safety check. The model learned exactly what it was asked to learn: which words tend to follow which words. Truth was never part of the objective, so it is not part of the result.

A real model is vastly better than this, because it conditions on far more context. That makes false continuations less likely, not impossible. The gap that produces hallucination is the same gap, narrowed.

The prompt shape that helps most

Everything else is downstream of this. Ground the answer, and give it a way out.

grounded_prompt.py
TEMPLATE = """Answer the question using ONLY the context below.
Quote the exact sentence you used, inside double quotes.
If the context does not contain the answer, reply exactly: NOT IN CONTEXT.

CONTEXT:
{context}

QUESTION: {question}
ANSWER:"""

def build(question, chunks):
    if not chunks:                       # nothing retrieved: do not call the model at all
        return None
    context = "\n".join(f"[{name}] {text}" for name, text in chunks)
    return TEMPLATE.format(context=context, question=question)

def answer(question, chunks):
    prompt = build(question, chunks)
    if prompt is None:
        return "NOT IN CONTEXT"          # refuse in our own code, before the model gets a turn
    return "<send this prompt to the model>\n" + prompt

print(answer("How much leave do I get?",
             [("policy", "Employees get 18 days of paid leave per year.")]))
print()
print("nothing retrieved ->", answer("Anything about pensions?", []))
Output
<send this prompt to the model>
Answer the question using ONLY the context below.
Quote the exact sentence you used, inside double quotes.
If the context does not contain the answer, reply exactly: NOT IN CONTEXT.

CONTEXT:
[policy] Employees get 18 days of paid leave per year.

QUESTION: How much leave do I get?
ANSWER:

nothing retrieved -> NOT IN CONTEXT

Three deliberate choices in that template.

A fixed refusal string. NOT IN CONTEXT is machine-checkable. "I'm not sure" is not, because the model will phrase it differently every time.

A required quote. It gives you something to verify mechanically, which the next script does.

Returning None when retrieval is empty. If nothing was found, calling the model at all invites it to answer from memory. Refuse in your own code, before the model gets a turn.

Verifying what came back

Never trust a quote without checking it. This takes fifteen lines.

verify_answer.py
import re

SOURCE = (
    "Leave policy, section 3. Every full-time employee gets 18 days of paid leave "
    "per year. Unused leave carries over up to a maximum of 10 days. Leave must be "
    "applied for at least 3 working days in advance."
)

# Pretend these came back from a model that was asked to quote its evidence.
ANSWER = (
    'You get 18 days of paid leave. The policy says "carries over up to a maximum of 10 days". '
    'Managers may grant an extra 5 days at their discretion, and the policy also says '
    '"leave may be encashed at the end of the year".'
)

def norm(t):
    return re.sub(r"\s+", " ", t.lower()).strip()

def check_quotes(answer, source):
    src = norm(source)
    return [(q, norm(q) in src) for q in re.findall(r'"([^"]+)"', answer)]

def check_numbers(answer, source):
    src_nums = set(re.findall(r"\d+", source))
    return [(n, n in src_nums) for n in re.findall(r"\d+", answer)]

print("quoted spans:")
for quote, ok in check_quotes(ANSWER, SOURCE):
    print(f"  [{'found' if ok else 'NOT IN SOURCE'}] \"{quote}\"")

print()
print("numbers used:")
for num, ok in check_numbers(ANSWER, SOURCE):
    print(f"  [{'found' if ok else 'NOT IN SOURCE'}] {num}")
Output
quoted spans:
  [found] "carries over up to a maximum of 10 days"
  [NOT IN SOURCE] "leave may be encashed at the end of the year"

numbers used:
  [found] 18
  [found] 10
  [NOT IN SOURCE] 5

Both fabrications were caught by string matching. No model, no embedding, no API call.

The quote check is the stronger of the two, because an invented quote is unambiguous evidence. The number check is a cheap smoke alarm: a number in the answer that appears nowhere in the source is worth a human look. Expect some false alarms from correct arithmetic over source numbers, and treat it as a flag rather than a verdict.

Catching the guesses by asking twice

When the model is recalling something solid, its answers stay put. When it is inventing, the details move. That difference is measurable.

self_consistency.py
import re
from collections import Counter

def key(answer):
    """Reduce an answer to the part worth comparing: its numbers and its names."""
    nums = re.findall(r"\d[\d,.]*", answer)
    caps = re.findall(r"\b[A-Z][a-z]+\b", answer)
    return (tuple(nums), tuple(sorted(set(caps))))

def agreement(answers):
    counts = Counter(key(a) for a in answers)
    top, n = counts.most_common(1)[0]
    return n / len(answers), top

# Five replies to the same question, from five separate calls at temperature 0.7.
RUN_A = [
    "The Kaveri river is about 800 km long.",
    "The Kaveri river is about 800 km long.",
    "The Kaveri is roughly 800 km in length.",
    "The Kaveri river is about 800 km long.",
    "The Kaveri river runs about 800 km.",
]
RUN_B = [
    "The report was published in 1997 by Menon.",
    "The report was published in 2001 by Menon.",
    "The report was published in 1994 by Sharma.",
    "The report was published in 2001 by Iyer.",
    "The report was published in 1997 by Menon.",
]

for name, answers in [("consistent question", RUN_A), ("shaky question", RUN_B)]:
    score, _ = agreement(answers)
    flag = "looks stable" if score >= 0.6 else "TREAT AS UNVERIFIED"
    print(f"{name:<20} agreement {score:.0%}   {flag}")
Output
consistent question  agreement 100%   looks stable
shaky question       agreement 40%   TREAT AS UNVERIFIED

The two answer lists are written out here so the example is reproducible. In production they come from five calls to the same model with temperature above zero.

Note carefully what this test does and does not do. High agreement means the model is confident, not that it is correct — a model can be consistently wrong about something it saw wrongly many times. Low agreement is the useful signal, and it is a reliable one. This is the practical core of the semantic-entropy work in the researcher tab.

Cost matters: five calls instead of one. Reserve it for claims where being wrong is expensive.

What does not work

Writing "do not hallucinate" in the prompt. It has no mechanism to obey. It cannot detect the state you are asking it to avoid.

Asking the model how confident it is. You get text that resembles a confidence statement. Studies repeatedly find these self-reported numbers poorly calibrated, and preference training makes them worse.

Assuming a bigger model fixes it. Bigger models hallucinate less on common facts and remain unreliable on rare ones. The tone often improves faster than the accuracy.

Assuming RAG fixes it. RAG helps enormously and does not eliminate it. The model can still ignore the context, blend it with memory, or answer confidently from a retrieved passage that was the wrong passage.

Common mistakes

Logging the answer but not the evidence. Without storing which chunks were retrieved, you cannot tell a retrieval failure from a generation failure. Log both, always.

Evaluating on questions you already know the answer to. Your test set gets quietly written to match what the model does well. Have someone else write it.

Treating a low-probability output as a hallucination detector. Token probability measures fluency, not truth. A confidently wrong answer often has higher probability than a hedged correct one.

Letting one wrong step ride. In a multi-step chain, an early invented fact is treated as established by every later step. Verify between steps, not only at the end.

Shipping a chatbot with no refusal path. If every input must produce an answer, you have guaranteed hallucination on every input outside its knowledge. Give the system somewhere to send "I do not know".

Try it yourself

Add a fourth fact to no_truth_check.py: put "the capital of italy is rome" into FACTS, then add italy and rome to the two loops at the bottom. You now get sixteen possible sentences, of which four are true. Three facts gave three true out of nine. Adding knowledge made the false fraction larger, not smaller.

Then take a topic you know deeply. Ask a model ten detailed questions about it and score the answers yourself. That number is the only honest estimate you will ever have of its reliability in fields you cannot check.

What to learn next

Researcher — Mathematics and papers.

Definitions worth keeping separate

  • Faithfulness failure: the output contradicts a provided source. Detectable by entailment checking against that source.
  • Factuality failure: the output contradicts the world. Requires an external oracle, and is therefore much harder to evaluate at scale.
  • Intrinsic: contradicts the input context. Extrinsic: unverifiable from the context, true or false.

Ji et al. (2023) established this taxonomy for generation broadly; Huang et al. (2023) survey the LLM-specific literature. Conflating faithfulness and factuality is the most common source of confused benchmark results, because a system can be perfectly faithful to a wrong retrieved document.

Why the objective produces it

Pretraining maximises Σ log p_θ(x_t | x_{<t}) over a text corpus. This is density estimation over strings. Nothing in the objective references a truth predicate, and no term distinguishes a true string from a fluent false one of equal corpus likelihood.

Two consequences follow directly, and both are formalised in recent work.

A calibration lower bound. Kalai and Vempala (2024), Calibrated Language Models Must Hallucinate, prove that a model calibrated to the training distribution must hallucinate at a rate bounded below by roughly the fraction of facts appearing exactly once in the corpus — the monofact rate, estimated by Good–Turing missing-mass arguments. For arbitrary facts with no learnable structure, the bound is not small. Crucially the result is architecture-independent: it applies to transformers, to n-gram models, and to anything else fitted to the same distribution. Reducing hallucination below this bound requires breaking calibration, which is what post-training and retrieval both do.

Evaluation rewards guessing. Kalai, Nachum, Vempala and Zhang (2025), Why Language Models Hallucinate, argue that persistence after pretraining is a scoring problem. Under binary-graded benchmarks, abstention scores zero and a guess has positive expected value, so a model optimised against those leaderboards learns to guess. Their prescription is a scoring change — explicit confidence targets with penalties for confident errors — rather than a new decoding trick. This reframes hallucination as partly a measurement artefact, which is a genuinely useful shift.

Aggravating factors

Snowballing. Zhang et al. (2023) show that once a model has committed to an incorrect claim, it produces further incorrect claims to remain consistent with it, including on questions it answers correctly in isolation. Verification must therefore be interleaved with generation, not deferred to the end.

Fine-tuning on new knowledge. Gekhman et al. (2024) find that SFT examples containing facts absent from pretraining are fitted slowly, and that as the model fits them its hallucination rate on unrelated questions rises. Teaching new facts through SFT trains the behaviour of asserting unsupported claims. This is the strongest available argument for retrieval over fine-tuning as a knowledge-injection mechanism.

RLHF degrades calibration. The GPT-4 system card reports a pretrained model well calibrated on MMLU, and the same model substantially miscalibrated after RLHF. Preference training optimises for answers humans like, and humans like confident answers.

Detection

Sampling-based consistency. SelfCheckGPT (Manakul et al., 2023) samples multiple responses and measures agreement, requiring no logits and no external resource. Effective and expensive.

Semantic entropy. Farquhar et al. (2024), in Nature, cluster sampled generations by bidirectional entailment before computing entropy, so paraphrases of one answer count once. This isolates uncertainty about meaning from uncertainty about wording, and outperforms naive token-entropy baselines across models and datasets. It is the strongest general-purpose detector currently published, and it detects confabulation — arbitrary, prompt-sensitive claims — rather than consistently held false beliefs.

Atomic fact scoring. FActScore (Min et al., 2023) decomposes a generation into atomic claims and verifies each against a knowledge source, giving a percentage rather than a binary judgement. The right granularity for long-form output, where a whole-answer label is nearly meaningless.

Internal-state probes. Linear probes on hidden activations detect statements the model itself represents as false at above-chance rates. Azaria and Mitchell (2023) and related work show the signal exists; generalisation across domains remains weak, and a probe that transfers reliably has not been demonstrated.

Mitigation, with honest effect sizes

  • Retrieval with attribution. The largest single reduction available, and the one with clear failure modes: retrieval miss, wrong passage, context ignored. See RAG.
  • Abstention training. R-Tuning (Zhang et al., 2024) fine-tunes on refusals for questions outside the model's parametric knowledge, improving willingness to decline without collapsing helpfulness.
  • Contrastive decoding. DoLa (Chuang et al., 2023) contrasts logits from later and earlier layers, exploiting the observation that factual content is refined in the upper layers. Cheap, no training, modest gains.
  • Activation steering. Inference-Time Intervention (Li et al., 2023) shifts activations along a truth-correlated direction found on TruthfulQA, improving that benchmark substantially. Transfer beyond the benchmark it was fitted on is the open question.
  • Constrained decoding and tool use. For anything with a checkable grammar — schemas, identifiers, arithmetic, dates — remove the freedom rather than correcting the output. This eliminates a whole class of failures instead of reducing it.

Benchmarks and their limits

TruthfulQA (Lin et al., 2022) targets imitative falsehoods, statements that are common in text and false. It is adversarially constructed, so scores do not estimate a real-world hallucination rate. HaluEval and FActScore cover recognition and long-form factuality respectively.

All published benchmarks now face contamination: they are on the public internet, and therefore in training corpora. Held-out evaluation on private data, refreshed over time, is the only defensible measurement. Report abstention rate alongside accuracy, or you are rewarding exactly the guessing behaviour Kalai et al. identify.

Papers

What to learn next

What to learn next

These follow on from what you just read.

  • LLM Development

    Hugging Face

    Hugging Face is a public library of ready-made AI models that you download with one line of Python instead of training your own.

  • LLM Development

    Ollama — run LLMs locally

    Ollama downloads a language model onto your own computer and runs it there, so nothing you type leaves your machine and nothing costs money per question.

  • LLM Development

    vLLM

    vLLM is a server that runs one language model on a GPU for many users at once, by never letting the graphics card sit idle waiting.