Question Answering

When two documents disagree

Retrieved chunks sometimes contradict each other, and a good system needs to detect that disagreement and resolve it deliberately, instead of quietly picking whichever fact came first.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Handling conflicting sources means noticing when two retrieved documents disagree, and deciding which one to trust.

Think about two relatives giving you different directions to the same wedding hall. One says turn left at the signal, the other says go straight. You trust whichever one has actually been there more recently.

Why it exists

A search step can retrieve more than one relevant document for a question, and those documents do not always agree. A company's leave policy from 2021 might say 18 days a year. The updated 2024 version might say 24.

A system that hands both to a model, with no guidance, leaves the disagreement for the model to sort out silently. Sometimes it picks the first one it reads. Sometimes it blends both numbers into something wrong either way.

The fix detects the conflict deliberately, then resolves it with a real rule. Usually that means recency, source authority, or both — never an arbitrary reading order.

How it works

Retrieved chunk 1: "Employees get 18 days of paid leave per year."
   source: hr-policy-2021.pdf,  last updated: 2021-04-01

Retrieved chunk 2: "Employees get 24 days of paid leave per year."
   source: hr-policy-2024.pdf,  last updated: 2024-01-15

Numbers found: 18 and 24  -> CONFLICT

Resolve by recency:
  2024 document is newer -> use "24 days"
  Flag the 2021 figure as outdated, do not silently discard it

Where you have already seen it

  • Wikipedia's "citation needed" and dispute flags, which surface disagreement instead of quietly picking one version.
  • News aggregators showing multiple outlets' reports on the same event side by side, when the details differ.
  • Product review summaries that show a rating spread, instead of hiding disagreement behind one averaged number.

Remember this

  • Retrieved documents can disagree, and a system needs to check for that, not assume agreement.
  • Recency and source authority are the two most common ways to resolve a conflict.
  • The losing source should be flagged for the user, not silently thrown away.

What to learn next

Developer — Code and libraries.

This detects a numeric conflict between two retrieved chunks and resolves it using the more recently updated source.

Setup

bash
python --version   # 3.9 or newer, only the standard library is used

Detecting and resolving a conflict

conflicting_sources.py
import re
from datetime import date

chunks = [
    {
        "text": "Employees get 18 days of paid leave per year.",
        "source": "hr-policy-2021.pdf",
        "last_updated": date(2021, 4, 1),
    },
    {
        "text": "Employees get 24 days of paid leave per year.",
        "source": "hr-policy-2024.pdf",
        "last_updated": date(2024, 1, 15),
    },
]

def extract_number(text):
    match = re.search(r"\d+", text)
    return match.group() if match else None

numbers = {extract_number(c["text"]) for c in chunks}
print("numbers found across retrieved chunks:", numbers)

if len(numbers) > 1:
    print("CONFLICT DETECTED: retrieved chunks disagree.")
    newest = max(chunks, key=lambda c: c["last_updated"])
    print(f"resolving by recency -> using {newest['source']} ({newest['last_updated']})")
    print("answer:", newest["text"])
    print("flagged for the user: an older source said something different")
    for c in chunks:
        if c is not newest:
            print(f"  - {c['source']} ({c['last_updated']}): {c['text']}")
else:
    print("no conflict, chunks agree.")
Output
numbers found across retrieved chunks: {'18', '24'}
CONFLICT DETECTED: retrieved chunks disagree.
resolving by recency -> using hr-policy-2024.pdf (2024-01-15)
answer: Employees get 24 days of paid leave per year.
flagged for the user: an older source said something different
  - hr-policy-2021.pdf (2021-04-01): Employees get 18 days of paid leave per year.

Line by line

extract_number() is a deliberately narrow conflict detector. It only catches numeric disagreements, using a plain regex over each chunk. Real conflicts are not always numeric — this is a starting point, not a general solution, and the researcher block covers more general approaches.

numbers = {extract_number(c["text"]) for c in chunks} builds a set. If every chunk agreed, this set would have exactly one entry. Two entries is the detection signal — simple, cheap, and specific to this narrow number-conflict case.

max(chunks, key=lambda c: c["last_updated"]) picks the chunk with the latest date, using Python's built-in date comparison. This is the entire resolution rule: newer wins.

The losing source is printed, not discarded. A system that resolves a conflict but hides that a conflict ever existed removes the user's ability to notice something changed, which matters more than it sounds — a leave-policy change is exactly the kind of update someone might want to know about directly.

Common mistakes

Assuming recency always means correctness. A newer document is not automatically more accurate — it could be a draft, or contain a typo introduced during an edit. Recency is a reasonable default rule, not a guarantee.

Detecting conflicts only on exact-duplicate topics. Two chunks phrased very differently, about the same underlying fact, can still disagree without sharing an easily-matched number or keyword. A regex-based detector like this one catches only the easy cases.

Silently averaging or blending conflicting numbers. "18 and 24 days, so about 21" is a plausible-sounding number that is also completely wrong — neither policy ever said 21 days. Never do this; always report a single authoritative source's number, alongside a stated resolution rule.

Ignoring source authority in favour of recency alone. An official policy PDF and a random internal chat message might have conflicting timestamps, but should not be weighted equally. Real systems often combine recency with a source-trust ranking, not recency alone.

Try it yourself

Add a third chunk with the same "24 days" figure as the 2024 document, but from a different, even newer source dated 2025.

Since max() compares by date, the newest entry should now win automatically, with no code changes — worth confirming directly, since it demonstrates the resolution rule generalises past exactly two sources.

What to learn next

Researcher — Mathematics and papers.

Conflict detection as a special case of consistency checking

The developer block's numeric-conflict detector is a narrow instance of a broader problem: given a set of retrieved passages {c_1, ..., c_k} about a shared topic, determine whether they are mutually consistent, and if not, which subset to trust. General-purpose detection requires more than regex matching on numbers — it requires determining whether two differently worded passages make the same, or different, factual claims, which is again the natural language inference problem covered in Natural language inference: treating one chunk as a premise and another as a hypothesis, and checking for contradiction rather than entailment.

Resolution strategies beyond recency

Recency-based resolution, as implemented in the developer block, assumes newer information supersedes older information — a reasonable default for policies, prices, and other facts that genuinely change over time, and a poor default for facts that do not change, where "newer" carries no informative signal at all.

Authority-based resolution ranks sources by a trust score — an official document outranking a forum post, for instance — independent of timestamp. This requires a maintained authority ranking over source types or specific documents, an operational cost the recency approach avoids.

Majority-vote resolution trusts whichever claim appears across the most independently retrieved sources. This is more robust to a single outdated or erroneous document, and fails when the majority is itself wrong, or when sources are not truly independent — several documents copied from one original source will vote together, inflating apparent consensus without adding real evidence.

Presenting the conflict rather than resolving it — showing the user both figures, with sources, and letting them decide — sidesteps the resolution problem by deferring it to the user. This is frequently the more honest choice for genuinely ambiguous or high-stakes conflicts, at the cost of shifting effort back onto the person asking.

Interaction with multi-hop reasoning

A conflict can also emerge across, rather than within, hops of a multi-hop question, covered earlier in this section. If hop 1 retrieves an outdated intermediate fact, the entire downstream chain inherits that error silently unless conflict detection runs at every hop, not only on the final answer. This compounding-error risk is one reason iterative multi-hop pipelines benefit from a consistency check after each retrieval step, not only once at the end.

Key references

Conflict detection and resolution in retrieval-augmented systems draws on the entailment and contradiction literature underlying natural language inference — see Natural language inference for the relevant benchmark and modelling references — combined with the general RAG survey framing in Gao et al. (2023), Retrieval-Augmented Generation for Large Language Models: A Survey (arXiv:2312.10997), which discusses source reliability as one dimension of retrieval quality.

Current state and open problems

General-purpose conflict detection — catching disagreements not reducible to a shared number or keyword — remains a harder, less mature problem than numeric conflict detection, and is usually implemented today with an LLM prompted to compare passages and flag contradictions, rather than a dedicated NLI model, largely because it needs to handle open-ended claim types with no fixed schema. There is no widely adopted, standardised evaluation benchmark specifically for RAG-context conflict detection at the time of writing, making it difficult to compare approaches rigorously across systems.

What to learn next