Data Leakage and Results That Are Too Good

Did the model already see your test set?

Large language models train on scraped internet text that often contains the benchmarks used to grade them — here is how contamination happens, how it is detected, and why every LLM score deserves a raised eyebrow.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

When a model trains on the whole internet, there is a real chance your exam questions were part of its homework.

Think of a student who prepared using every guidebook in the market. The examiner picks questions from a well-known question bank — which, it turns out, was reprinted inside two of those guidebooks. The student answers "from memory" without anyone intending to cheat. The exam, the student and the guidebooks are all behaving normally. The combination is broken.

Large language models are that student. Their training data is scraped web text at enormous scale. Benchmark questions — exams for AI — live on the web too: in papers, blog posts, tutorial sites, discussion forums. The overlap is called benchmark contamination.

Why it exists

Benchmarks are published openly so everyone can compare models fairly. Publication puts them on the internet. Training scrapes the internet. Nobody has to make a mistake for the exam to end up inside the homework — the pipeline does it automatically.

And once people discuss a benchmark online — quoting questions, explaining answers — the contamination multiplies. The better known the exam, the more copies of its answer key float around the web.

How it works

benchmark published ──→ appears on websites, papers, forums
                                   │
internet scraped for training ─────┘
                                   ▼
              model memorises questions + answers
                                   ▼
benchmark run on model ──→ score partly measures MEMORY
                            and gets reported as ABILITY

The score is not fake — the model really does produce right answers. What breaks is the meaning: the number no longer predicts how the model handles genuinely new problems.

A real example you have seen

Demos where a chatbot solves a famous riddle instantly. Change one detail — the colours, the names, the numbers — and it confidently gives the original answer, now wrong. It recalled the riddle rather than reasoning through it. That trick — rephrase and watch the score change — is also one of the serious detection tools.

Remember this

  • Public benchmark + scraped training data = accidental contamination, no villain required.
  • A contaminated score measures memory but gets sold as ability.
  • Distrust famous-benchmark scores; trust held-back or freshly written tests.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

The checker below is pure Python; numpy is only for the ecosystem's sake. Verified on Python 3.10, CPU. Deterministic — your output should match exactly.

An n-gram overlap checker

The standard first-line contamination audit asks: does any long word-sequence from the benchmark appear inside the training text? An n-gram is a run of n consecutive words; long runs (8–13 words) almost never repeat across unrelated texts, so a shared one is strong evidence of copying.

contamination_check.py
def ngrams(text, n=8):
    words = text.lower().replace(",", "").replace(".", "").split()
    return {" ".join(words[i:i + n]) for i in range(len(words) - n + 1)}

# stand-in for one scraped web page in a training corpus
train_page = ("Quiz night answers. The capital of Australia is Canberra, "
              "not Sydney, which surprises many people every single time.")

benchmark = [
    "The capital of Australia is Canberra, not Sydney, which surprises many people.",
    "Which planet in our solar system has the most confirmed moons?",
]

train_grams = ngrams(train_page)
for question in benchmark:
    overlap = ngrams(question) & train_grams
    verdict = "CONTAMINATED" if overlap else "looks clean"
    print(f"{verdict:12s} | {question[:52]}")
    if overlap:
        print(f"             | shared 8-gram: '{sorted(overlap)[0]}'")
Output
CONTAMINATED | The capital of Australia is Canberra, not Sydney, wh
             | shared 8-gram: 'australia is canberra not sydney which surprises many'
looks clean  | Which planet in our solar system has the most confir

The first question shares a long run of words with the scraped page — copied, beyond reasonable doubt. The second shares nothing that long. At real scale this exact logic runs over terabytes with hashed n-grams and bloom filters — the GPT-3 team used 13-gram overlap for their audit. The principle fits in fifteen lines.

What the checker misses

Exact n-grams catch verbatim copies only. The harder cases:

  • Paraphrases: a translated or reworded question shares no 8-gram with its source. Embedding similarity between benchmark items and training passages catches some of this.
  • Answer-only contamination: the model saw the answer discussed without the exact question.
  • Closed training data: for API models you cannot scan the corpus at all. Detection must become behavioural — see the researcher block.

Working defensively as a practitioner

  • Prefer tiny, private evals for your own decisions. Fifty questions written by your team, never posted anywhere, beat a famous public benchmark for judging models on your task. Agent evaluation builds this habit for agent systems.
  • Perturb before you trust. Rename entities, reorder options, change numbers. A large score drop on minor rephrasing is memory confessing.
  • Date-partition when possible. Evaluate on material created after the model's training cutoff — new exams, new problems, new news.
  • Never publish your private eval set — the moment it is indexed, its clock starts ticking.

Common mistakes

Comparing models with different training cutoffs on one benchmark. The newer model may have had more chances to absorb the benchmark. Score gaps partly measure scrape dates.

Fine-tuning on data that contains your eval. The contamination pipeline in miniature, and entirely under your control to prevent: n-gram-scan your fine-tuning set against your eval set before training. See fine-tuning.

Treating "the provider deduplicated" as a guarantee. Published audits routinely find residual overlap after dedup; filters have bugs (famously, GPT-3's own contamination filter had one).

Writing eval questions by asking the model to generate them. The model generates what it has seen. You are laundering its memory into your exam.

Try it yourself

Lower n from 8 to 3 and rerun. The clean question now shows "contamination" — from generic phrases like "of australia is". Push n to 12 and the real copy escapes too. Sit with the trade-off you have discovered: false alarms versus misses, controlled by one integer.

What to learn next

Researcher — Mathematics and papers.

Documented contamination in the wild

  • Brown et al. (2020), Language Models are Few-Shot Learners (GPT-3), Appendix C: 13-gram overlap audit across all benchmarks; a filtering bug shipped anyway; per-benchmark "dirty" vs "clean" splits showed mostly modest deltas — the honest self-audit that set the genre's template.
  • Dodge et al. (2021), Documenting Large Webtext Corpora: A Case Study on C4, EMNLP: found test sets of multiple standard NLP benchmarks verbatim inside C4.
  • Lee et al. (2022), Deduplicating Training Data Makes Language Models Better, ACL: near-duplicate mass in web corpora and its memorisation consequences; suffix-array exact matching plus MinHash — the scaled version of this lesson's checker.
  • Sainz et al. (2023), NLP Evaluation in Trouble, EMNLP Findings: community-maintained LM Contamination Index; documents ChatGPT reciting benchmark examples on request.
  • Zhang et al. (2024), A Careful Examination of LLM Benchmark Contamination (GSM1k): freshly written GSM8k-style problems; several models dropped sharply relative to their GSM8k scores — the perturbation test at publication scale.

Detection without corpus access

For closed models, detection becomes statistical inference over model behaviour:

  • Canonical-order tests — Oren et al. (2023), Proving Test Set Contamination in Black Box Language Models, ICLR 2024: benchmark datasets are exchangeable (order-free); a model assigning systematically higher likelihood to the published ordering of examples than to random permutations proves the ordering was trained on. An elegant hypothesis test with exact p-values.
  • Guided completion — Golchin and Surdeanu (2023), Time Travel in LLMs, ICLR 2024: prompt with dataset name and a partial example; near-verbatim completion of the remainder indicates memorisation, scored against reference-free baselines.
  • Membership-inference adjacent: perplexity gaps between benchmark items and matched fresh items; the general MIA literature (Carlini et al., 2021, Extracting Training Data from Large Language Models, USENIX Security) supplies the machinery and its limits.
  • Canary strings: BIG-bench (Srivastava et al., 2022) embeds a unique GUID so any model emitting it proves benchmark ingestion — contamination detection designed into the benchmark.

Structural responses

Static public benchmarks depreciate; the field's countermeasures trade comparability for validity:

  • Refreshing benchmarks: LiveBench (White et al., 2024) sources questions from post-cutoff material monthly; LiveCodeBench does the same for code.
  • Held-out private sets administered by third parties (e.g., SEAL-style private leaderboards; Kaggle private splits — the test-set overfitting defence at ecosystem scale).
  • Human preference evaluation (LMSYS Chatbot Arena, Chiang et al., 2024): fresh prompts from live users cannot pre-exist in training data, at the cost of measuring "preference" rather than task accuracy.
  • Perturbation families: benchmark items as templates with resampled surface forms; performance is reported as a distribution over instantiations.

The equilibrium is unstable by construction: any benchmark that matters gets discussed, and any discussion gets scraped. Treat every static public number as a ceiling on ability, dated by the model's cutoff.

What to learn next