Data Leakage and Results That Are Too Good
Duplicate rows across your splits
When the same row lands in both training and test data, your model gets marked on questions it has memorised — dedup before you split, every time.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
If the same row appears in both your training data and your test data, the test stops measuring learning and starts measuring memory.
Imagine preparing for an exam using a practice book. Unknown to you, the examiner copied ten questions straight from that book onto the final paper. You answer those ten from memory, word for word. Your mark goes up, but your understanding did not.
Duplicated rows do the same to a model. The test set is supposed to be an unseen exam. A duplicate is a question the model has already seen, with the answer attached.
Why it exists
Real datasets are assembled, merged, appended and re-exported. Duplicates creep in through ordinary, boring accidents:
- The same file ingested twice on different days.
- A join that matched one record to several.
- The same customer signing up twice with the same details.
- Two data sources that partly overlap.
Then you split randomly into training and test sets. A random split scatters each duplicate pair randomly — so around half of the pairs end up with one copy on each side. Nobody decided to cheat. The cheat assembled itself.
How it works
dataset with duplicates random split
row A ──┐ train: A, B, C
row B │ shuffle ──→ test : B*, D
row B* │ ↑
row C │ B* has a twin in training —
row D ──┘ the model answers it from memoryThe more your model memorises (big trees, big networks), the more those twin rows inflate the score.
A real example you have seen
Song recommendation datasets famously contain the same track many times — as an album version, a remaster, and a compilation entry. A model tested on "unseen" songs is often re-scoring songs it trained on, under a different ID. Researchers found the same problem inside famous photo datasets used to rank AI models for years.
Remember this
- Duplicates plus a random split means the test set overlaps the training set.
- The inflation grows with how much your model can memorise.
- Remove duplicates before splitting, and check for near-copies, not only exact ones.
What to learn next
- Fitting the scaler before splitting — the third suspect: preparation that peeked.
- Data cleaning — where deduplication lives in a real pipeline.
- The same person in train and test — duplicates at the level of people rather than rows.
Developer — Code and libraries.
Setup
pip install scikit-learn pandasVerified with scikit-learn 1.7.2, pandas 2.2.3, numpy 1.26.4, CPU. Seeded, so your numbers should match.
Coin-flip labels, 80% accuracy
The labels below are random coin flips. There is nothing to learn, so any honest score is near 0.5. Watch what duplicates do.
import numpy as np
import pandas as pd
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import train_test_split
rng = np.random.default_rng(2)
base = pd.DataFrame(rng.normal(size=(300, 5)).round(2), columns=list("abcde"))
base["label"] = rng.integers(0, 2, 300) # labels are coin flips: nothing to learn
# 150 rows got ingested twice, a very ordinary data bug
df = pd.concat([base, base.sample(150, random_state=2)]).reset_index(drop=True)
print("duplicate rows:", df.duplicated().sum())
Xtr, Xte, ytr, yte = train_test_split(df[list("abcde")], df["label"], random_state=2)
tree = DecisionTreeClassifier(random_state=2).fit(Xtr, ytr)
print("accuracy with duplicates:", round(tree.score(Xte, yte), 3))
clean = df.drop_duplicates()
Xtr, Xte, ytr, yte = train_test_split(clean[list("abcde")], clean["label"], random_state=2)
tree = DecisionTreeClassifier(random_state=2).fit(Xtr, ytr)
print("accuracy after dedup :", round(tree.score(Xte, yte), 3))duplicate rows: 150 accuracy with duplicates: 0.796 accuracy after dedup : 0.48
Eighty percent accuracy on pure noise. After deduplication, the truth: coin-flip performance.
The walkthrough
Why a decision tree? An unlimited-depth tree memorises its training data perfectly, which makes it the ideal instrument for demonstrating contamination. Any model that memorises — which includes every large neural network — inflates the same way, only less visibly.
The two lines that matter are df.duplicated().sum() before splitting, and drop_duplicates() before splitting. Order is everything: deduplicating after the split cannot reunite twins that already straddle it.
Why 0.796 and not 1.0? Only test rows whose twin landed in training get answered from memory. The rest are genuine unseen noise, scored at coin-flip rates. Real contamination is partial, which makes it look plausible — a 0.796 raises fewer eyebrows than a 1.0.
Near-duplicates: the harder version
Exact-match dedup misses rows that differ by a timestamp, a whitespace, or one resized pixel. Check for near-duplicates — rows identical in the columns that matter. Append this to the same file, so df is already in scope:
key_columns = list("abcde") # identity columns, ignoring ids and timestamps
near = df.duplicated(subset=key_columns).sum()
print("duplicates ignoring non-key columns:", near)duplicates ignoring non-key columns: 150
For text and images, near-duplicate detection needs fuzzier tools — hashing shortened text, or comparing embeddings (text embeddings covers the representation side).
Common mistakes
Deduplicating after splitting. The twins have already crossed. Dedup is a pre-split operation, always.
Checking only exact duplicates. A re-exported row with a new updated_at timestamp defeats duplicated() on full rows. Dedup on the columns that define identity.
Forgetting that IDs hide duplicates. Two rows with different customer_id values can still be the same person. If identity matters, that is group leakage, the same disease at a different scale.
Assuming fresh data cannot overlap old data. When you retrain next quarter, last quarter's test rows are sitting in the new training pull. Version your splits — data versioning exists for exactly this.
Try it yourself
Change DecisionTreeClassifier to LogisticRegression. Predict what happens to the contaminated score before running. Then explain the result using the memorisation argument above.
What to learn next
- Fitting the scaler before splitting — the third suspect: preparation that peeked.
- Data cleaning — where deduplication lives in a real pipeline.
- The same person in train and test — duplicates at the level of people rather than rows.
Researcher — Mathematics and papers.
The inflation, quantified
Let a fraction $c$ of test rows have an exact twin in training, and let the model interpolate its training set. The expected measured accuracy is
$$ \hat{A} = c \cdot 1 + (1 - c) \cdot A $$
Where:
- $c$ — the contaminated fraction of the test set.
- $A$ — true accuracy on genuinely unseen data.
- $\hat{A}$ — the inflated estimate you observe.
In the demo, 150 of 450 rows are copies. A test row that is one of the 300 pair-members has its twin in training with probability equal to the training fraction, 0.75. Working through the expectation gives $\hat{A} \approx 0.75$ for $A = 0.5$ — close to the observed 0.796, the gap being finite-sample noise. The formula also shows why contamination is vicious for strong models: as models approach interpolation, the first term hits its ceiling of 1.
Documented contamination in benchmark datasets
- Barz and Denzler (2020), Do We Train on Test Data? Purging CIFAR of Near-Duplicates, Journal of Imaging: roughly 3% of CIFAR-10 and 10% of CIFAR-100 test images have near-duplicates in their training sets. They release corrected test sets (ciFAIR); model rankings shift measurably.
- Recht et al. (2019), Do ImageNet Classifiers Generalize to ImageNet?, ICML: building truly fresh test sets drops absolute accuracy for every model, part of which reflects the ways original test sets have been absorbed by the community.
- Lee et al. (2022), Deduplicating Training Data Makes Language Models Better, ACL: web-scale corpora contain enormous near-duplicate mass; dedup reduces memorised emissions roughly tenfold and improves perplexity at equal compute. Their two tools — exact substring matching via suffix arrays, and MinHash for approximate document matching — are the standard heavy machinery.
Near-duplicate detection at scale
For $n$ documents, pairwise comparison is $O(n^2)$ and infeasible. MinHash with locality-sensitive hashing (Broder, 1997, On the resemblance and containment of documents) estimates Jaccard similarity of shingle sets in sub-quadratic time: documents agreeing on enough hash bands become candidate pairs. For images, perceptual hashes or embedding-space nearest neighbours serve the same role, with the similarity threshold as the policy decision — too tight misses paraphrases, too loose deletes legitimate diversity.
The LLM-scale version of this problem — benchmark answers dissolved into pretraining text — has its own lesson: did the model already see your test set?
Split hygiene as versioned state
Contamination recurs because splits are recomputed while data accumulates. The durable fix is treating the test set as an immutable, versioned artefact: freeze row identities (hashes, not indices), store the manifest alongside the data version, and assert disjointness in CI on every retrain. Gorman and Bedrick (2019), We Need to Talk about Standard Splits, ACL, argue the complementary point — that single frozen splits carry their own fragility, so report across resplits with dedup enforced each time.
What to learn next
- Fitting the scaler before splitting — the third suspect: preparation that peeked.
- Data cleaning — where deduplication lives in a real pipeline.
- The same person in train and test — duplicates at the level of people rather than rows.