Reading and Reimplementing Papers
A staged plan for reimplementing a paper
Reimplement in stages that each end in a check you can pass or fail — pick one number to reproduce, get the evaluation working before the method, and memorise eight examples before touching the full dataset.
- 11 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Reimplement a paper in small stages, where each stage ends in a check that passes or fails outright — and never start with the interesting part.
Nobody learns to drive on a highway. You start in an empty ground with nothing to hit. Then a quiet lane at eight in the morning. Then a main road. Each stage adds one new difficulty, so when something goes wrong you know which difficulty caused it.
Most failed reimplementations skip straight to the highway. Full dataset, full model, full method, first run. Then the number comes out wrong and there are two hundred possible reasons.
Decide what "done" means, first
Before writing any code, pick one number from the paper that you are trying to reproduce.
Not "reimplement the paper". One number. A single cell of a single table, with the dataset, the metric and the setting written down.
This sounds like a small thing. It is the difference between a project that ends and a project that drifts for three months. Without a target, every result is ambiguous — you can always tell yourself it is close enough, or that the remaining gap is somebody else's fault.
Write the target down before you start, in a file, with the page number.
The stages
stage 0 pick ONE number to reproduce → written down?
stage 1 evaluation working, dumb model → does the metric run?
stage 2 method transcribed, checked → matches a reference?
stage 3 memorise a handful of examples → loss near zero?
stage 4 small run, trusted baseline beside → in the same neighbourhood?
stage 5 full run, compare to the paper → hit the target?Each arrow is a question with a yes or no answer. If a stage fails, you fix it there. You never carry a failing stage forward, because a broken foundation makes every later result meaningless.
Why evaluation comes before the method
Stage 1 builds the scoring machinery with a deliberately stupid model — one that predicts the most common answer, or a random answer.
This feels backwards, and it is the most valuable stage in the list. It tells you the data loads, the metric runs, and what the score looks like when nothing has been learned. Every later number gets compared against that floor.
Teams that skip it routinely celebrate a result that turns out to be worse than guessing.
The honest bit
Stage 3 — memorise a handful of examples — is the one people find silly and the one that catches the most bugs.
Take eight examples. Train on those eight only. The model should learn them almost perfectly, because memorising eight things is easy. If it cannot, something is broken in the plumbing: labels misaligned, gradients not flowing, a learning rate a thousand times too small. The learning rate is the size of each training step. Nothing at full scale will work either, and at full scale you will not be able to tell.
Remember this
- Pick one number to reproduce, and write it down before coding.
- Build the evaluation before the method, with a stupid model.
- Memorise eight examples before touching the full dataset.
What to learn next
- Why your reimplementation is three points worse — closing the last gap.
- Overfit one batch — stage 3 as a permanent habit.
- Ablation studies — the discipline that keeps late-stage debugging interpretable.
Developer — Code and libraries.
Setup
pip install scikit-learnVerified with scikit-learn 1.7.2 and numpy 1.26.4 on CPU. The whole script runs in a few seconds and downloads nothing — load_digits ships inside scikit-learn.
The worked target
We continue with Kingma and Ba (2015). Section 6.1 of the paper trains multi-class logistic regression on MNIST with Adam. Our scaled-down target: train the same model with a hand-written Adam and land near a trusted implementation of the same model.
Be clear about the scaling down. load_digits is 1,797 images of 8×8 pixels — a different, far easier dataset than MNIST's 70,000 images of 28×28. We are reproducing the shape of the experiment on a laptop, not its numbers. Saying so out loud is part of doing this honestly.
Stages 3, 4 and 5 in one script
import numpy as np
from sklearn.datasets import load_digits
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
def softmax_xent(W, b, X, Y):
"""Loss and its gradients for multi-class logistic regression."""
z = X @ W + b
z = z - z.max(axis=1, keepdims=True) # stable log-sum-exp
logp = z - np.log(np.exp(z).sum(axis=1, keepdims=True))
loss = -(Y * logp).mean(axis=0).sum()
gz = (np.exp(logp) - Y) / len(X)
return loss, X.T @ gz, gz.sum(axis=0)
def fit(X, Y, steps, alpha=0.01, seed=0):
rng = np.random.default_rng(seed)
W = rng.normal(0, 0.01, (X.shape[1], Y.shape[1]))
b = np.zeros(Y.shape[1])
mW, vW = np.zeros_like(W), np.zeros_like(W)
mb, vb = np.zeros_like(b), np.zeros_like(b)
for t in range(1, steps + 1):
loss, gW, gb = softmax_xent(W, b, X, Y)
for p, g, m, v in ((W, gW, mW, vW), (b, gb, mb, vb)):
m *= 0.9 # m and v persist across steps
m += 0.1 * g
v *= 0.999
v += 0.001 * g * g
p -= alpha * (m / (1 - 0.9**t)) / (np.sqrt(v / (1 - 0.999**t)) + 1e-8)
return W, b, loss
d = load_digits()
X, Y = d.data / 16.0, np.eye(10)[d.target]
Xtr, Xte, Ytr, Yte = train_test_split(X, Y, test_size=0.3, random_state=0)
# stage 3 -- can it memorise eight examples? if not, the update is broken.
_, _, small = fit(Xtr[:8], Ytr[:8], steps=500)
print(f"stage 3 loss on 8 examples after 500 steps: {small:.6f}")
# stage 4 -- the real (small) dataset
W, b, loss = fit(Xtr, Ytr, steps=2000)
acc = (np.argmax(Xte @ W + b, axis=1) == np.argmax(Yte, axis=1)).mean()
print(f"stage 4 training loss {loss:.4f} test accuracy {acc:.4f}")
# stage 5 -- against a trusted implementation of the same model
ref = LogisticRegression(max_iter=5000).fit(Xtr, np.argmax(Ytr, axis=1))
print(f"stage 5 sklearn LogisticRegression test accuracy "
f"{ref.score(Xte, np.argmax(Yte, axis=1)):.4f}")stage 3 loss on 8 examples after 500 steps: 0.003988 stage 4 training loss 0.0167 test accuracy 0.9574 stage 5 sklearn LogisticRegression test accuracy 0.9611
Read the three lines as three verdicts.
Stage 3 passed. Loss near zero on eight examples means gradients flow, labels line up with predictions, and the optimiser moves parameters in the right direction. Everything downstream is now worth debugging.
Stage 4 produced a number, which is meaningless on its own.
Stage 5 gave it meaning. 95.74% against 96.11% from a mature, well-tested implementation of the same model. A gap of 0.37 points on the same data is the neighbourhood you want — close enough that the method is right, far enough that something differs. Chasing that remainder is the next lesson.
The walkthrough
Stage 3 costs six seconds and rules out half of all bugs. A model that cannot memorise eight examples has a plumbing fault, not a tuning problem. This is overfitting one batch, applied before there is anything to overfit.
The optimiser is the transcribed Adam from the previous lesson, now updating two parameter arrays rather than one. m and v are mutated in place inside the tuple loop, so their state survives between steps. That in-place update is easy to break during a refactor — assert m.any() after the first step if you ever suspect it.
softmax_xent subtracts the row maximum before exponentiating. Straight transcription of the written formula overflows on large scores; this rewrite is algebraically identical and stable.
The reference is the same model, not the same paper. scikit-learn's LogisticRegression uses a different solver and adds a small L2 penalty by default. It is a sanity anchor, not a ground truth. Comparing against a different model would tell you nothing.
The reported loss comes from the last forward pass, before the final update. At 2,000 steps that distinction is invisible; at 5 steps it is not.
Budget the project before starting it
Before stage 1, write down three numbers: how long one full run takes, how many full runs you can afford, and what a full run costs in money. A paper that trained for 500 GPU-days cannot be reproduced end to end on a laptop, and knowing that on day one changes the project into something achievable — one ablation, one dataset, one figure.
Then scale down deliberately and write down what you changed: smaller dataset, fewer layers, shorter schedule. A scaled-down reproduction that states its scaling is a real result. An unstated one is a broken result.
Common mistakes
Starting with the interesting part. The novel mechanism is the fun bit and the wrong bit to write first. Without evaluation, you cannot tell whether it is working.
No baseline in the same script. A number with nothing beside it cannot be interpreted. Keep a trusted reference and a dumb baseline running in every experiment, forever.
Fixing stage 4 by tuning. When the small run underperforms, the reflex is to raise the learning rate or add epochs. If stage 3 was skipped, you may be tuning around a broken gradient. Go back rather than forward.
Changing several things per run. Once results are close, change one thing at a time and record what happened. See ablation studies.
Not writing the target down. Without stage 0 in a file, the project ends when you get bored rather than when you are done.
Try it yourself
Break stage 3 on purpose: shuffle Ytr so labels no longer match their images, and rerun. The loss on eight examples will still fall near zero — memorising noise is easy — while stage 4 collapses to roughly 10%. That contrast is exactly why stage 3 checks plumbing and stage 4 checks learning, and why you need both.
What to learn next
- Why your reimplementation is three points worse — closing the last gap.
- Overfit one batch — stage 3 as a permanent habit.
- Ablation studies — the discipline that keeps late-stage debugging interpretable.
Researcher — Mathematics and papers.
Three words that mean different things to different committees
"Reproducible" is used in opposite senses across communities. ACM's artifact badging separates repeatability (same team, same setup), reproducibility (a different team, the original artifacts) and replicability (a different team, independently built artifacts) — and the assignment of the latter two labels was itself revised in 2020 to align with NISO usage. Metrology and the social sciences use the pair the other way round.
The practical consequence: never write "we reproduced the paper" without saying which of the three you did. Running the authors' released code on the authors' released data is a far weaker claim than building an independent implementation from the text and reaching the same number, and only the second tells you the paper is correct.
Structuring the attempt so it is reviewable
- Freeze the target in a file. Table, row, column, dataset, metric, and the exact number, committed before any results exist. This is the pre-registration idea imported from clinical trials, and it removes the temptation to redefine success after seeing the output.
- Separate the three failure classes. A gap arises from an implementation error, an unstated detail, or a claim that does not hold. Design the stages so each is ruled out in turn: numerical equivalence rules out the first, ablations over plausible detail choices bound the second, and only then is the third even discussable.
- Version the environment. Library versions, random seeds, hardware and dtype all move results. See config files, not arguments and random seeds and reproducibility.
- Log every run, including the failures. The failed runs contain the evidence about which details mattered, and they are the part of a reproduction report that other people cannot obtain cheaply.
What community-scale reproduction efforts found
The ML Reproducibility Challenge, running annually since the 2018 ICLR edition, assigns published papers to independent teams who attempt reimplementation and publish reports. Two consistent findings across editions: most reproduction failures trace to unstated hyperparameters and preprocessing rather than to false claims; and the reports that are useful are the ones that state precisely how far they scaled down.
On the venue side, Pineau et al. (2021), Improving Reproducibility in Machine Learning Research (JMLR), document the NeurIPS code-submission policy and reproducibility checklist and their measured effect on code availability. The checklist itself is a serviceable outline for what your own reproduction report should contain.
Deciding when to stop
Set the stopping rule with the target, before results exist. Reasonable rules: within one standard deviation of the reported number over five seeds; or an ablation-explained gap where every remaining difference is attributed to a stated deviation. An open-ended "get closer" has no terminating condition and consumes arbitrary time.
If the gap will not close, the reproduction report is still a contribution — provided it documents the search over unstated details rather than concluding from a single configuration. That search is the subject of finding the details the paper left out.
What to learn next
- Why your reimplementation is three points worse — closing the last gap.
- Overfit one batch — stage 3 as a permanent habit.
- Ablation studies — the discipline that keeps late-stage debugging interpretable.