Reading and Reimplementing Papers

Reading a results table sceptically

The bold number in a results table is the most carefully engineered object in the paper — here is the checklist that tells you whether it means anything.

On this page 6
  1. Why scepticism is the right default
  2. The seven questions
  3. How it works
  4. A real example you have seen
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A results table shows the comparison the authors chose to show, run the way they chose to run it — read it as an argument, not as a measurement.

Two coaching classes advertise on the same wall. One says "98% success rate". The other says "our topper scored 99 marks". Neither is lying. One counted everybody who passed anything; the other found a single student. The poster is designed, and design is where the meaning went.

A results table is a poster. Every row, column and bold number in it was a decision.

Why scepticism is the right default

Authors are not villains. They are people under pressure whose paper is rejected if the new method does not win.

That pressure has predictable effects. You tune your own method for weeks and the baseline — the method you are comparing against — for an afternoon. You report the datasets where it worked. You run once, see a good number, and stop. None of this requires dishonesty — it happens to careful people who mean well.

Whole research areas have been checked this way afterwards, and the checks are humbling. One 2019 study looked at eighteen neural recommendation papers from top conferences. Only seven could be reproduced at all. Six of those seven were then beaten by simple, decades-old methods, given a fair amount of tuning.

The seven questions

Run these against any table before believing it.

  1. Same data, same split? Different splits are different exams.
  2. Was the baseline tuned as hard as the new method? This is where most gains come from.
  3. Are there error bars? A single run is a single dice roll.
  4. Is the gain bigger than the wobble? If repeated runs vary by 2 and the gain is 1, there is no gain.
  5. How many datasets, and were any left out? Winning on 3 of 12 is a different story from winning on 3 of 3.
  6. Was the test set used to pick the model? The test set is the data held back for the final score. Once it picks the model, it is not a test set.
  7. Where did the baseline numbers come from? Copied from an older paper is not the same as re-run here.

How it works

   what the table shows          what you want to know
   ────────────────────          ─────────────────────
   ours     94.2  (bold)         is 94.2 - 93.1 bigger than
   baseline 93.1                 the wobble between reruns?

                                 was the baseline given the
                                 same tuning budget?

                                 how many datasets did this
                                 NOT happen on?

A real example you have seen

Product benchmark charts follow the same rules. A phone launch shows a bar chart against a rival, on the one test where it wins, with the axis starting at 80 rather than 0. You already read those sceptically. Research tables deserve the same reflex, and get it far less often.

Remember this

  • A results table is an argument, built from choices you cannot see.
  • The most common source of a "gain" is an under-tuned baseline.
  • No error bars means no result — one run tells you almost nothing.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Verified with scikit-learn 1.7.2 and numpy 1.26.4 on CPU. Each script runs in a few seconds.

How big is the wobble?

Before you can judge a gap, you need to know the noise floor. Here is the same model, on the same data, with the same split — changing only the random seed.

seed_wobble.py
import numpy as np
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.neural_network import MLPClassifier

X, y = make_classification(n_samples=2000, n_features=20, n_informative=8,
                           flip_y=0.15, random_state=0)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=0)

scores = []
for seed in range(10):                      # only the seed changes, nothing else
    model = MLPClassifier(hidden_layer_sizes=(32,), early_stopping=True,
                          max_iter=500, random_state=seed)
    scores.append(model.fit(Xtr, ytr).score(Xte, yte))

scores = np.array(scores)
print("ten seeds, same model, same split, same data:")
print("  ", np.round(scores, 4))
print(f"  best {scores.max():.4f}   worst {scores.min():.4f}   spread {scores.ptp():.4f}")
print(f"  mean {scores.mean():.4f}   std   {scores.std(ddof=1):.4f}")
Output
ten seeds, same model, same split, same data:
   [0.8317 0.8117 0.85   0.835  0.8217 0.85   0.8317 0.84   0.8317 0.8317]
  best 0.8500   worst 0.8117   spread 0.0383
  mean 0.8335   std   0.0116

Read that carefully. This is one model. Reporting 85.0 for it and 81.2 for it, in two rows of a table, would be a fabricated 3.8-point improvement — and every number would be real.

Any paper claiming a gain of one or two points on a setup like this, from single runs, has told you nothing.

The gap has a distribution too

Now a genuine improvement, measured ten times over different splits.

gap_across_splits.py
import numpy as np
from sklearn.datasets import make_classification
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

X, y = make_classification(n_samples=3000, n_features=20, n_informative=8,
                           flip_y=0.15, random_state=0)

gaps = []
for split in range(10):                       # only the train/test split changes
    Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=split)
    fancy = HistGradientBoostingClassifier(random_state=0).fit(Xtr, ytr)
    base = make_pipeline(StandardScaler(),
                         LogisticRegression(max_iter=2000)).fit(Xtr, ytr)
    gaps.append(fancy.score(Xte, yte) - base.score(Xte, yte))

gaps = np.array(gaps)
print("new method minus baseline, per split:")
print("  ", np.round(gaps, 4))
print(f"  mean gap {gaps.mean():+.4f}   std {gaps.std(ddof=1):.4f}")
print(f"  splits where the baseline won: {(gaps < 0).sum()} of 10")
print(f"  best single split you could report: {gaps.max():+.4f}")
Output
new method minus baseline, per split:
   [0.0078 0.0078 0.0089 0.0211 0.0122 0.01   0.0089 0.0133 0.0133 0.0156]
  mean gap +0.0119   std 0.0042
  splits where the baseline won: 0 of 10
  best single split you could report: +0.0211

This gain is real. It survives every split, which is the strongest evidence in this lesson.

And the honest number is +1.19 points, while the reportable number is +2.11 points. Picking the friendly split nearly doubles the headline without a single fabricated measurement. That is the mechanism behind most inflated tables — no misconduct required, only a choice about which run to write down.

The walkthrough

Seed variance and split variance are different beasts. The first script holds the data fixed and varies initialisation. The second holds the model fixed and varies the data. A paper should report both; most report neither. See seed variance and error bars for how to present them.

early_stopping=True makes the runs converge cleanly, so the spread you see is real seed sensitivity rather than an artefact of stopping at an arbitrary iteration.

Ten repeats is a floor, not a target. With ten samples, the standard error on the standard deviation is itself large. For a real claim, use enough repeats that adding more stops moving the mean.

"Baseline never won" is worth more than the mean gap. Ten out of ten is a sign test with a p-value near 0.001, computed without assuming anything about the distribution of the differences.

Common mistakes

Comparing your mean against their single number. Baseline numbers copied from an earlier paper were produced under a different protocol, split and preprocessing pipeline. Re-run the baseline yourself, or say plainly that you did not.

Treating a leaderboard position as a measurement. Public leaderboards are optimised against by hundreds of teams, which is test-set overfitting at industrial scale. Rank movements of a few tenths carry no information.

Reporting the best run. Reporting max over seeds guarantees a number your readers cannot reproduce. Report mean and spread, and say how many runs.

Tuning your method and not the baseline. Give both the same search budget, from the same search space, and say what the budget was. If that feels like it will hurt your result, that feeling is the finding.

Reading bold as "significantly better". Bold usually marks "highest number in the column". It rarely survives a statistical test, and almost never accounts for multiple comparisons across a wide table.

Try it yourself

Change flip_y=0.15 to flip_y=0.30 in the second script and rerun. Predict first: does the mean gap shrink, and does the baseline start winning some splits? Then reduce n_samples to 600 and watch how much easier it becomes to report whatever conclusion you would like.

What to learn next

Researcher — Mathematics and papers.

Field-scale audits, and what they found

The single-paper failures are not the interesting part. What matters is that when whole subfields were re-evaluated under fair protocols, the reported progress often did not survive.

  • Recommender systems. Dacrema, Cremonesi and Jannach (2019), Are We Really Making Much Progress? (RecSys, best paper): of 18 neural recommendation methods from top venues, 7 were reproducible with reasonable effort, and 6 of those 7 were outperformed by well-tuned nearest-neighbour and graph baselines.
  • Language models. Melis, Dyer and Blunsom (2018), On the State of the Art of Evaluation in Neural Language Models (ICLR): with a shared, large hyperparameter search, standard LSTM architectures matched or beat the newer architectures that had been reported to supersede them on Penn Treebank and WikiText-2.
  • Deep metric learning. Musgrave, Belongie and Lim (2020), A Metric Learning Reality Check (ECCV): improvements claimed over roughly a decade largely vanished under a fair protocol; common flaws included using the test set for model selection and comparing against baselines with weaker backbones.
  • GANs. Lucic et al. (2018), Are GANs Created Equal? (NeurIPS): with enough hyperparameter search and random restarts, most GAN variants reached comparable FID, and no variant consistently beat the original non-saturating loss.
  • Optimisers. Schmidt, Schneider and Hennig (2021), Descending through a Crowded Valley (ICML): a large benchmark of roughly fifteen optimisers over eight tasks found no dominant method, and found that tuning one well-known optimiser was competitive with trying many.

The recurring mechanism is unequal tuning budgets, not fraud. Comparing method A tuned for 200 trials against method B tuned for 5 measures the search, not the method.

Where the variance comes from

Bouthillier et al. (2021), Accounting for Variance in Machine Learning Benchmarks, decompose performance variance into sources: weight initialisation, data order, data augmentation randomness, dropout masks, split choice, and hyperparameter search randomness. Two findings shape practice. First, the total is often larger than the reported improvements in the literature. Second, varying only the seed understates it — data-split variance usually dominates, so the honest protocol randomises the split as well and reports the standard deviation across the full pipeline.

Testing a difference properly

  • Multiple datasets — Demšar (2006), Statistical Comparisons of Classifiers over Multiple Data Sets (JMLR): use the Wilcoxon signed-ranks test for two methods, and Friedman with a post-hoc Nemenyi test for several. Averaging accuracy across datasets is meaningless, since the scales differ.
  • One dataset, resampled — Dietterich (1998) shows the naive resampled paired t-test has badly inflated Type I error because the training sets overlap; the $5\times2$ cross-validated paired t-test is the standard repair.
  • Bayesian alternative — Benavoli et al. (2017) argue for posterior probabilities over a region of practical equivalence, which answers "is the difference large enough to care about?" rather than "is it non-zero?"
  • Multiplicity. A table with 12 datasets and 8 methods contains hundreds of implicit comparisons. Without correction, some bold entries are guaranteed by chance.

Benchmark decay

Even a clean protocol degrades over time as a community optimises against a fixed test set. Recht et al. (2019), Do ImageNet Classifiers Generalize to ImageNet?, built new test sets by replicating the original collection process: accuracy dropped by roughly 11 to 14 points across models, while the ranking was largely preserved. The practical reading is that absolute benchmark numbers age badly, relative ordering ages better, and small differences on a heavily used benchmark are the least trustworthy quantity in the table.

A minimum reporting standard

For your own tables: report mean and standard deviation over at least five runs varying seed and split; state the tuning budget and search space for every method including baselines; report every dataset you ran, including losses; name the exact metric implementation; and say whether baseline numbers were re-run or cited. Papers that do this are noticeably harder to publish and noticeably more useful.

What to learn next