Data Leakage and Results That Are Too Good

Duplicate rows across your splits

When the same row lands in both training and test data, your model gets marked on questions it has memorised — dedup before you split, every time.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

If the same row appears in both your training data and your test data, the test stops measuring learning and starts measuring memory.

Imagine preparing for an exam using a practice book. Unknown to you, the examiner copied ten questions straight from that book onto the final paper. You answer those ten from memory, word for word. Your mark goes up, but your understanding did not.

Duplicated rows do the same to a model. The test set is supposed to be an unseen exam. A duplicate is a question the model has already seen, with the answer attached.

Why it exists

Real datasets are assembled, merged, appended and re-exported. Duplicates creep in through ordinary, boring accidents:

  • The same file ingested twice on different days.
  • A join that matched one record to several.
  • The same customer signing up twice with the same details.
  • Two data sources that partly overlap.

Then you split randomly into training and test sets. A random split scatters each duplicate pair randomly — so around half of the pairs end up with one copy on each side. Nobody decided to cheat. The cheat assembled itself.

How it works

dataset with duplicates          random split

  row A ──┐                    train: A, B, C
  row B   │  shuffle  ──→      test : B*, D
  row B*  │                          ↑
  row C   │              B* has a twin in training —
  row D ──┘              the model answers it from memory

The more your model memorises (big trees, big networks), the more those twin rows inflate the score.

A real example you have seen

Song recommendation datasets famously contain the same track many times — as an album version, a remaster, and a compilation entry. A model tested on "unseen" songs is often re-scoring songs it trained on, under a different ID. Researchers found the same problem inside famous photo datasets used to rank AI models for years.

Remember this

  • Duplicates plus a random split means the test set overlaps the training set.
  • The inflation grows with how much your model can memorise.
  • Remove duplicates before splitting, and check for near-copies, not only exact ones.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn pandas

Verified with scikit-learn 1.7.2, pandas 2.2.3, numpy 1.26.4, CPU. Seeded, so your numbers should match.

Coin-flip labels, 80% accuracy

The labels below are random coin flips. There is nothing to learn, so any honest score is near 0.5. Watch what duplicates do.

duplicate_inflation.py
import numpy as np
import pandas as pd
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(2)
base = pd.DataFrame(rng.normal(size=(300, 5)).round(2), columns=list("abcde"))
base["label"] = rng.integers(0, 2, 300)   # labels are coin flips: nothing to learn

# 150 rows got ingested twice, a very ordinary data bug
df = pd.concat([base, base.sample(150, random_state=2)]).reset_index(drop=True)
print("duplicate rows:", df.duplicated().sum())

Xtr, Xte, ytr, yte = train_test_split(df[list("abcde")], df["label"], random_state=2)
tree = DecisionTreeClassifier(random_state=2).fit(Xtr, ytr)
print("accuracy with duplicates:", round(tree.score(Xte, yte), 3))

clean = df.drop_duplicates()
Xtr, Xte, ytr, yte = train_test_split(clean[list("abcde")], clean["label"], random_state=2)
tree = DecisionTreeClassifier(random_state=2).fit(Xtr, ytr)
print("accuracy after dedup    :", round(tree.score(Xte, yte), 3))
Output
duplicate rows: 150
accuracy with duplicates: 0.796
accuracy after dedup    : 0.48

Eighty percent accuracy on pure noise. After deduplication, the truth: coin-flip performance.

The walkthrough

Why a decision tree? An unlimited-depth tree memorises its training data perfectly, which makes it the ideal instrument for demonstrating contamination. Any model that memorises — which includes every large neural network — inflates the same way, only less visibly.

The two lines that matter are df.duplicated().sum() before splitting, and drop_duplicates() before splitting. Order is everything: deduplicating after the split cannot reunite twins that already straddle it.

Why 0.796 and not 1.0? Only test rows whose twin landed in training get answered from memory. The rest are genuine unseen noise, scored at coin-flip rates. Real contamination is partial, which makes it look plausible — a 0.796 raises fewer eyebrows than a 1.0.

Near-duplicates: the harder version

Exact-match dedup misses rows that differ by a timestamp, a whitespace, or one resized pixel. Check for near-duplicates — rows identical in the columns that matter. Append this to the same file, so df is already in scope:

duplicate_inflation.py (continued)
key_columns = list("abcde")            # identity columns, ignoring ids and timestamps
near = df.duplicated(subset=key_columns).sum()
print("duplicates ignoring non-key columns:", near)
Output
duplicates ignoring non-key columns: 150

For text and images, near-duplicate detection needs fuzzier tools — hashing shortened text, or comparing embeddings (text embeddings covers the representation side).

Common mistakes

Deduplicating after splitting. The twins have already crossed. Dedup is a pre-split operation, always.

Checking only exact duplicates. A re-exported row with a new updated_at timestamp defeats duplicated() on full rows. Dedup on the columns that define identity.

Forgetting that IDs hide duplicates. Two rows with different customer_id values can still be the same person. If identity matters, that is group leakage, the same disease at a different scale.

Assuming fresh data cannot overlap old data. When you retrain next quarter, last quarter's test rows are sitting in the new training pull. Version your splits — data versioning exists for exactly this.

Try it yourself

Change DecisionTreeClassifier to LogisticRegression. Predict what happens to the contaminated score before running. Then explain the result using the memorisation argument above.

What to learn next

Researcher — Mathematics and papers.

The inflation, quantified

Let a fraction $c$ of test rows have an exact twin in training, and let the model interpolate its training set. The expected measured accuracy is

$$ \hat{A} = c \cdot 1 + (1 - c) \cdot A $$

Where:

  • $c$ — the contaminated fraction of the test set.
  • $A$ — true accuracy on genuinely unseen data.
  • $\hat{A}$ — the inflated estimate you observe.

In the demo, 150 of 450 rows are copies. A test row that is one of the 300 pair-members has its twin in training with probability equal to the training fraction, 0.75. Working through the expectation gives $\hat{A} \approx 0.75$ for $A = 0.5$ — close to the observed 0.796, the gap being finite-sample noise. The formula also shows why contamination is vicious for strong models: as models approach interpolation, the first term hits its ceiling of 1.

Documented contamination in benchmark datasets

  • Barz and Denzler (2020), Do We Train on Test Data? Purging CIFAR of Near-Duplicates, Journal of Imaging: roughly 3% of CIFAR-10 and 10% of CIFAR-100 test images have near-duplicates in their training sets. They release corrected test sets (ciFAIR); model rankings shift measurably.
  • Recht et al. (2019), Do ImageNet Classifiers Generalize to ImageNet?, ICML: building truly fresh test sets drops absolute accuracy for every model, part of which reflects the ways original test sets have been absorbed by the community.
  • Lee et al. (2022), Deduplicating Training Data Makes Language Models Better, ACL: web-scale corpora contain enormous near-duplicate mass; dedup reduces memorised emissions roughly tenfold and improves perplexity at equal compute. Their two tools — exact substring matching via suffix arrays, and MinHash for approximate document matching — are the standard heavy machinery.

Near-duplicate detection at scale

For $n$ documents, pairwise comparison is $O(n^2)$ and infeasible. MinHash with locality-sensitive hashing (Broder, 1997, On the resemblance and containment of documents) estimates Jaccard similarity of shingle sets in sub-quadratic time: documents agreeing on enough hash bands become candidate pairs. For images, perceptual hashes or embedding-space nearest neighbours serve the same role, with the similarity threshold as the policy decision — too tight misses paraphrases, too loose deletes legitimate diversity.

The LLM-scale version of this problem — benchmark answers dissolved into pretraining text — has its own lesson: did the model already see your test set?

Split hygiene as versioned state

Contamination recurs because splits are recomputed while data accumulates. The durable fix is treating the test set as an immutable, versioned artefact: freeze row identities (hashes, not indices), store the manifest alongside the data version, and assert disjointness in CI on every retrain. Gorman and Bedrick (2019), We Need to Talk about Standard Splits, ACL, argue the complementary point — that single frozen splits carry their own fragility, so report across resplits with dedup enforced each time.

What to learn next