Reading and Reimplementing Papers
Reading a results table sceptically
The bold number in a results table is the most carefully engineered object in the paper — here is the checklist that tells you whether it means anything.
- 11 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A results table shows the comparison the authors chose to show, run the way they chose to run it — read it as an argument, not as a measurement.
Two coaching classes advertise on the same wall. One says "98% success rate". The other says "our topper scored 99 marks". Neither is lying. One counted everybody who passed anything; the other found a single student. The poster is designed, and design is where the meaning went.
A results table is a poster. Every row, column and bold number in it was a decision.
Why scepticism is the right default
Authors are not villains. They are people under pressure whose paper is rejected if the new method does not win.
That pressure has predictable effects. You tune your own method for weeks and the baseline — the method you are comparing against — for an afternoon. You report the datasets where it worked. You run once, see a good number, and stop. None of this requires dishonesty — it happens to careful people who mean well.
Whole research areas have been checked this way afterwards, and the checks are humbling. One 2019 study looked at eighteen neural recommendation papers from top conferences. Only seven could be reproduced at all. Six of those seven were then beaten by simple, decades-old methods, given a fair amount of tuning.
The seven questions
Run these against any table before believing it.
- Same data, same split? Different splits are different exams.
- Was the baseline tuned as hard as the new method? This is where most gains come from.
- Are there error bars? A single run is a single dice roll.
- Is the gain bigger than the wobble? If repeated runs vary by 2 and the gain is 1, there is no gain.
- How many datasets, and were any left out? Winning on 3 of 12 is a different story from winning on 3 of 3.
- Was the test set used to pick the model? The test set is the data held back for the final score. Once it picks the model, it is not a test set.
- Where did the baseline numbers come from? Copied from an older paper is not the same as re-run here.
How it works
what the table shows what you want to know
──────────────────── ─────────────────────
ours 94.2 (bold) is 94.2 - 93.1 bigger than
baseline 93.1 the wobble between reruns?
was the baseline given the
same tuning budget?
how many datasets did this
NOT happen on?A real example you have seen
Product benchmark charts follow the same rules. A phone launch shows a bar chart against a rival, on the one test where it wins, with the axis starting at 80 rather than 0. You already read those sceptically. Research tables deserve the same reflex, and get it far less often.
Remember this
- A results table is an argument, built from choices you cannot see.
- The most common source of a "gain" is an under-tuned baseline.
- No error bars means no result — one run tells you almost nothing.
What to learn next
- Seed variance and error bars — how to measure and present the wobble.
- Baselines you must beat — the row that is missing from most tables.
- Test-set overfitting — how a benchmark quietly stops being a test.
Developer — Code and libraries.
Setup
pip install scikit-learnVerified with scikit-learn 1.7.2 and numpy 1.26.4 on CPU. Each script runs in a few seconds.
How big is the wobble?
Before you can judge a gap, you need to know the noise floor. Here is the same model, on the same data, with the same split — changing only the random seed.
import numpy as np
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.neural_network import MLPClassifier
X, y = make_classification(n_samples=2000, n_features=20, n_informative=8,
flip_y=0.15, random_state=0)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=0)
scores = []
for seed in range(10): # only the seed changes, nothing else
model = MLPClassifier(hidden_layer_sizes=(32,), early_stopping=True,
max_iter=500, random_state=seed)
scores.append(model.fit(Xtr, ytr).score(Xte, yte))
scores = np.array(scores)
print("ten seeds, same model, same split, same data:")
print(" ", np.round(scores, 4))
print(f" best {scores.max():.4f} worst {scores.min():.4f} spread {scores.ptp():.4f}")
print(f" mean {scores.mean():.4f} std {scores.std(ddof=1):.4f}")ten seeds, same model, same split, same data: [0.8317 0.8117 0.85 0.835 0.8217 0.85 0.8317 0.84 0.8317 0.8317] best 0.8500 worst 0.8117 spread 0.0383 mean 0.8335 std 0.0116
Read that carefully. This is one model. Reporting 85.0 for it and 81.2 for it, in two rows of a table, would be a fabricated 3.8-point improvement — and every number would be real.
Any paper claiming a gain of one or two points on a setup like this, from single runs, has told you nothing.
The gap has a distribution too
Now a genuine improvement, measured ten times over different splits.
import numpy as np
from sklearn.datasets import make_classification
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
X, y = make_classification(n_samples=3000, n_features=20, n_informative=8,
flip_y=0.15, random_state=0)
gaps = []
for split in range(10): # only the train/test split changes
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=split)
fancy = HistGradientBoostingClassifier(random_state=0).fit(Xtr, ytr)
base = make_pipeline(StandardScaler(),
LogisticRegression(max_iter=2000)).fit(Xtr, ytr)
gaps.append(fancy.score(Xte, yte) - base.score(Xte, yte))
gaps = np.array(gaps)
print("new method minus baseline, per split:")
print(" ", np.round(gaps, 4))
print(f" mean gap {gaps.mean():+.4f} std {gaps.std(ddof=1):.4f}")
print(f" splits where the baseline won: {(gaps < 0).sum()} of 10")
print(f" best single split you could report: {gaps.max():+.4f}")new method minus baseline, per split: [0.0078 0.0078 0.0089 0.0211 0.0122 0.01 0.0089 0.0133 0.0133 0.0156] mean gap +0.0119 std 0.0042 splits where the baseline won: 0 of 10 best single split you could report: +0.0211
This gain is real. It survives every split, which is the strongest evidence in this lesson.
And the honest number is +1.19 points, while the reportable number is +2.11 points. Picking the friendly split nearly doubles the headline without a single fabricated measurement. That is the mechanism behind most inflated tables — no misconduct required, only a choice about which run to write down.
The walkthrough
Seed variance and split variance are different beasts. The first script holds the data fixed and varies initialisation. The second holds the model fixed and varies the data. A paper should report both; most report neither. See seed variance and error bars for how to present them.
early_stopping=True makes the runs converge cleanly, so the spread you see is real seed sensitivity rather than an artefact of stopping at an arbitrary iteration.
Ten repeats is a floor, not a target. With ten samples, the standard error on the standard deviation is itself large. For a real claim, use enough repeats that adding more stops moving the mean.
"Baseline never won" is worth more than the mean gap. Ten out of ten is a sign test with a p-value near 0.001, computed without assuming anything about the distribution of the differences.
Common mistakes
Comparing your mean against their single number. Baseline numbers copied from an earlier paper were produced under a different protocol, split and preprocessing pipeline. Re-run the baseline yourself, or say plainly that you did not.
Treating a leaderboard position as a measurement. Public leaderboards are optimised against by hundreds of teams, which is test-set overfitting at industrial scale. Rank movements of a few tenths carry no information.
Reporting the best run. Reporting max over seeds guarantees a number your readers cannot reproduce. Report mean and spread, and say how many runs.
Tuning your method and not the baseline. Give both the same search budget, from the same search space, and say what the budget was. If that feels like it will hurt your result, that feeling is the finding.
Reading bold as "significantly better". Bold usually marks "highest number in the column". It rarely survives a statistical test, and almost never accounts for multiple comparisons across a wide table.
Try it yourself
Change flip_y=0.15 to flip_y=0.30 in the second script and rerun. Predict first: does the mean gap shrink, and does the baseline start winning some splits? Then reduce n_samples to 600 and watch how much easier it becomes to report whatever conclusion you would like.
What to learn next
- Seed variance and error bars — how to measure and present the wobble.
- Baselines you must beat — the row that is missing from most tables.
- Test-set overfitting — how a benchmark quietly stops being a test.
Researcher — Mathematics and papers.
Field-scale audits, and what they found
The single-paper failures are not the interesting part. What matters is that when whole subfields were re-evaluated under fair protocols, the reported progress often did not survive.
- Recommender systems. Dacrema, Cremonesi and Jannach (2019), Are We Really Making Much Progress? (RecSys, best paper): of 18 neural recommendation methods from top venues, 7 were reproducible with reasonable effort, and 6 of those 7 were outperformed by well-tuned nearest-neighbour and graph baselines.
- Language models. Melis, Dyer and Blunsom (2018), On the State of the Art of Evaluation in Neural Language Models (ICLR): with a shared, large hyperparameter search, standard LSTM architectures matched or beat the newer architectures that had been reported to supersede them on Penn Treebank and WikiText-2.
- Deep metric learning. Musgrave, Belongie and Lim (2020), A Metric Learning Reality Check (ECCV): improvements claimed over roughly a decade largely vanished under a fair protocol; common flaws included using the test set for model selection and comparing against baselines with weaker backbones.
- GANs. Lucic et al. (2018), Are GANs Created Equal? (NeurIPS): with enough hyperparameter search and random restarts, most GAN variants reached comparable FID, and no variant consistently beat the original non-saturating loss.
- Optimisers. Schmidt, Schneider and Hennig (2021), Descending through a Crowded Valley (ICML): a large benchmark of roughly fifteen optimisers over eight tasks found no dominant method, and found that tuning one well-known optimiser was competitive with trying many.
The recurring mechanism is unequal tuning budgets, not fraud. Comparing method A tuned for 200 trials against method B tuned for 5 measures the search, not the method.
Where the variance comes from
Bouthillier et al. (2021), Accounting for Variance in Machine Learning Benchmarks, decompose performance variance into sources: weight initialisation, data order, data augmentation randomness, dropout masks, split choice, and hyperparameter search randomness. Two findings shape practice. First, the total is often larger than the reported improvements in the literature. Second, varying only the seed understates it — data-split variance usually dominates, so the honest protocol randomises the split as well and reports the standard deviation across the full pipeline.
Testing a difference properly
- Multiple datasets — Demšar (2006), Statistical Comparisons of Classifiers over Multiple Data Sets (JMLR): use the Wilcoxon signed-ranks test for two methods, and Friedman with a post-hoc Nemenyi test for several. Averaging accuracy across datasets is meaningless, since the scales differ.
- One dataset, resampled — Dietterich (1998) shows the naive resampled paired t-test has badly inflated Type I error because the training sets overlap; the $5\times2$ cross-validated paired t-test is the standard repair.
- Bayesian alternative — Benavoli et al. (2017) argue for posterior probabilities over a region of practical equivalence, which answers "is the difference large enough to care about?" rather than "is it non-zero?"
- Multiplicity. A table with 12 datasets and 8 methods contains hundreds of implicit comparisons. Without correction, some bold entries are guaranteed by chance.
Benchmark decay
Even a clean protocol degrades over time as a community optimises against a fixed test set. Recht et al. (2019), Do ImageNet Classifiers Generalize to ImageNet?, built new test sets by replicating the original collection process: accuracy dropped by roughly 11 to 14 points across models, while the ranking was largely preserved. The practical reading is that absolute benchmark numbers age badly, relative ordering ages better, and small differences on a heavily used benchmark are the least trustworthy quantity in the table.
A minimum reporting standard
For your own tables: report mean and standard deviation over at least five runs varying seed and split; state the tuning budget and search space for every method including baselines; report every dataset you ran, including losses; name the exact metric implementation; and say whether baseline numbers were re-run or cited. Papers that do this are noticeably harder to publish and noticeably more useful.
What to learn next
- Seed variance and error bars — how to measure and present the wobble.
- Baselines you must beat — the row that is missing from most tables.
- Test-set overfitting — how a benchmark quietly stops being a test.