Question Answering

Checking that an answer really came from the source

A generated answer can cite a real source and still say something the source never claimed, so attribution checking verifies the words in the answer are actually supported by the words in the context.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Verifying attribution means checking that a generated answer is actually backed up by the source it claims to come from.

Think about a school essay where every claim needs a page number from the textbook next to it. A teacher checking the essay flips to that page. If the page says something different, the citation was fake, even though a page number was given.

Why it exists

A retrieval system can hand a model the exact right passage. The model can still write an answer that goes beyond, or contradicts, what that passage says. A citation being present does not prove it is accurate.

This matters because a wrong answer with a source attached looks more trustworthy than a wrong answer with none. That makes an unchecked, wrongly-attributed answer more dangerous, not less.

Attribution checking closes that gap. After an answer is generated, a separate check compares its claims against the retrieved text, flagging anything unsupported.

How it works

Retrieved context:
  "Infosys reported quarterly revenue of 38,000 crore rupees, a rise
   of 6 percent. The company also announced a buyback worth 9,300 crore."

Generated answer A:
  "Infosys revenue grew 6 percent to 38,000 crore rupees."
  -> CHECK: every claim matches the context. Supported.

Generated answer B:
  "Infosys revenue grew 12 percent, its fastest pace in five years."
  -> CHECK: "12 percent" and "fastest pace" are not in the context.
  -> NOT SUPPORTED. Flag before showing to the user.

Where you have already seen it

  • News fact-checking sidebars, which compare a claim against the article or document it is supposedly sourced from.
  • Academic plagiarism and citation checkers, which verify a citation actually says what the citing text claims.
  • AI tools with a "verify" or "check sources" button, which re-check the model's own output against what it retrieved.

Remember this

  • A citation being shown does not prove the citation is accurate.
  • Attribution checking compares the specific claims in an answer against the specific text of the source.
  • This step runs after generation, as a check, not as part of writing the answer itself.

What to learn next

Developer — Code and libraries.

This checks how many of an answer's content words are actually present in the retrieved context, flagging words that appear to be unsupported.

Setup

bash
python --version   # 3.9 or newer, only the standard library is used

A simple lexical support check

attribution_check.py
import re

STOPWORDS = {"the", "a", "an", "is", "was", "were", "of", "in", "to", "and",
             "or", "on", "at", "for", "by", "with", "that", "this", "it",
             "its", "as", "from"}

def content_words(text):
    words = re.findall(r"[a-z0-9]+", text.lower())
    return [w for w in words if w not in STOPWORDS]

def support_score(answer, context):
    ans_words = content_words(answer)
    ctx_words = set(content_words(context))
    if not ans_words:
        return 0.0
    supported = [w for w in ans_words if w in ctx_words]
    return len(supported) / len(ans_words)

context = (
    "Infosys reported quarterly revenue of 38,000 crore rupees, a rise of 6 "
    "percent year on year. The company also announced a share buyback worth "
    "9,300 crore rupees."
)

answer_supported = "Infosys revenue grew 6 percent to 38,000 crore rupees."
answer_hallucinated = "Infosys revenue grew 12 percent, its fastest pace in five years."

for label, ans in [("supported answer", answer_supported),
                    ("hallucinated answer", answer_hallucinated)]:
    score = support_score(ans, context)
    ctx_words = set(content_words(context))
    unsupported = [w for w in content_words(ans) if w not in ctx_words]
    print(f"{label}: {ans!r}")
    print(f"  support score: {score:.2f}")
    print(f"  words not found in context: {unsupported}")
Output
supported answer: 'Infosys revenue grew 6 percent to 38,000 crore rupees.'
  support score: 0.89
  words not found in context: ['grew']
hallucinated answer: 'Infosys revenue grew 12 percent, its fastest pace in five years.'
  support score: 0.33
  words not found in context: ['grew', '12', 'fastest', 'pace', 'five', 'years']

Line by line

content_words() strips out common function words, so the comparison focuses on the words carrying actual meaning — numbers, names, and content-bearing terms — rather than being diluted by "the," "a," and "is."

The supported answer scores 0.89, not 1.00. The word "grew" is flagged as unsupported, even though the context plainly implies growth. This is the honest limitation of lexical matching: the context says "a rise of," not "grew," and this simple check has no way to know those mean the same thing.

The hallucinated answer scores 0.33, and every fabricated detail shows up explicitly in the unsupported-words list: "12," "fastest," "pace," "five," "years" — none of them appear anywhere in the context, because the model invented them.

This is a heuristic, not a certainty. A low score is a strong signal something may be unsupported. It is not proof, and a high score is not proof of correctness either — see the researcher block for a harder, more general check.

Common mistakes

Treating this as a hallucination detector for meaning, rather than wording alone. As shown above, "grew" versus "a rise of" triggers a false flag. This method catches wording that has no basis in the source at all far better than it distinguishes a valid paraphrase from a genuine fabrication.

Setting one universal support-score threshold across very different answer types. A one-word factual answer and a three-sentence explanation behave very differently under this scoring. Tune thresholds per answer type, or use this as one signal among several, not a single hard cutoff.

Running this only on user-facing answers, never during development. This check is cheap enough to run on every generated answer during testing, catching a prompt or retrieval regression before it reaches a real user.

Assuming a perfect score means the answer is fully correct. An answer that copies every word from the context verbatim scores 1.00 by this method, even if it copies the wrong part of the context, out of context, or misrepresents what the source actually concluded.

Try it yourself

Write a third answer that paraphrases the context correctly but uses entirely different words — for example, "Infosys grew its top line by 6 percent last quarter and returned cash to shareholders through a buyback." Run it through support_score.

Expect a lower score than answer_supported, despite being accurate — "top line," "returned cash," and "shareholders" are not literal words from the context. This is the exact limitation flagged above, made concrete with your own example.

What to learn next

Researcher — Mathematics and papers.

Faithfulness versus correctness

The RAG evaluation literature separates two properties that are easy to conflate. Faithfulness (or attribution, or groundedness) asks whether an answer's claims are supported by the retrieved context, regardless of whether that context is itself true. Correctness asks whether the answer is true, regardless of what the retrieved context said. A system can be faithful to a wrong or outdated source and be unfaithful while still, by coincidence, producing a correct answer. This lesson addresses faithfulness specifically — it is a check on the generation step, not a check on the retrieval or the underlying facts.

From lexical overlap to entailment

The developer block's word-overlap heuristic is a crude approximation of a better-defined problem: textual entailment, whether a context passage entails an answer claim, covered in Natural language inference. A production faithfulness checker typically replaces lexical overlap with a natural language inference model, or a general-purpose LLM prompted to judge entailment, scoring whether each claim in the answer is entailed by, contradicted by, or unrelated to the retrieved context. This resolves the paraphrase problem demonstrated in the developer block — "grew" versus "a rise of" — since entailment reasons over meaning, not surface word overlap.

Automated faithfulness metrics

Es, James, Espinosa-Anke & Schockaert (2023), Ragas: Automated Evaluation of Retrieval Augmented Generation (arXiv:2309.15217), formalises a faithfulness metric along these lines: decompose the generated answer into individual claims, then check each claim against the retrieved context using an LLM-based judge, and report the fraction of claims that are supported. This claim-decomposition approach is more robust than checking the answer as one undivided block, since a single long answer often mixes supported and unsupported claims together, and a whole-answer score would obscure exactly which part failed.

Cost and reliability trade-offs

Lexical overlap is free, fast, and fully deterministic, at the cost of the paraphrase blind spot shown in the developer block. NLI-model-based checking costs one model call per claim, is more semantically aware, and inherits whatever calibration issues the specific NLI model carries — see Natural language inference for its own limitations. LLM-as-judge checking, as in Ragas, is the most flexible and the most expensive, and introduces its own reliability question: an LLM judge can itself be miscalibrated or inconsistent, covered in LLM-as-a-judge. No method here is free of trade-offs; production systems typically pick based on the acceptable cost per check, not a single method being strictly best.

Key references

  • Es, S., James, J., Espinosa-Anke, L. & Schockaert, S. (2023). Ragas: Automated Evaluation of Retrieval Augmented Generation. arXiv:2309.15217

Current state and open problems

No attribution-checking method here is fully reliable on its own. Lexical methods miss valid paraphrases; NLI models are trained on specific entailment datasets that may not transfer to a given domain; LLM judges carry their own, less well-understood failure modes. Combining multiple signals — lexical overlap as a cheap first pass, an NLI or LLM check on flagged cases — is a common practical compromise, trading some cost for better coverage than any single method alone. A fully reliable, cheap, general-purpose faithfulness checker remains an open problem in the RAG evaluation literature.

What to learn next