Reading and Reimplementing Papers

Why your reimplementation is three points worse

The gap between your number and the paper's is almost never the method — it is a schedule, a normalisation constant, an evaluation protocol, or one word in a config that means two different things.

On this page 6
  1. The order to check things in
  2. Why evaluation comes first
  3. How it works
  4. The honest bit
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

When your version scores worse than the paper, the cause is nearly always a small unstated detail — not the method, and not your understanding of it.

You cook your aunt's rajma from her written recipe and it comes out flat. The recipe is not wrong. She soaks the beans overnight, her cooker takes three whistles not five, and she adds the tomatoes only after the onions have gone properly brown. None of that is written down, because to her it is not a step. It is how cooking works.

Papers are written by people with the same blind spot.

The order to check things in

Sort your suspects by how often they are the answer.

  1. Evaluation protocol. Are you measuring the same thing on the same data?
  2. Data preparation. Different scaling, different splits, different cleaning.
  3. Schedule. How the step size changes during training, and for how long.
  4. Regularisation settings. Applied to everything, or to some parts only.
  5. Which run got reported. Best of ten looks better than the average of ten.
  6. The method itself. Last, because it is rarely the answer.

Most people start at 6 and work upward. That is why reproductions take months.

Why evaluation comes first

Because a mismatch there makes every other comparison meaningless, and it is common.

Papers report scores in ways that quietly differ. Best score seen during training, or the score at the end. The average over ten runs, or the best. On a validation set, the data used while tuning, or a test set, the data saved for the final answer. With the model's weights from the final step, or a smoothed average of recent steps. Each of these choices moves the number by a point or two, and papers often name none of them.

Before hunting anything else, make sure you and the paper are grading the same exam the same way.

How it works

   your number   93.1
   paper's       96.4
                 ────
   gap            3.3   ← this is a LIST, not a single cause

   evaluation protocol   1.8
   schedule difference   0.9
   data normalisation    0.4
   seed noise            0.2
   remaining             0.0   ← the goal

The work is to turn one unexplained gap into a list of explained ones. A remaining gap of zero is rare. A remaining gap you can name is the real target.

The honest bit

Sometimes the gap does not close, and the reason is that the reported number was fortunate. A single run reported without error bars can sit at the top of its own distribution. Your average of five runs is then correctly lower.

This is not a failure of your work. Establishing it takes real effort, and reporting it is a genuine contribution.

Remember this

  • The gap is a list of small causes, not one big one.
  • Check the evaluation protocol first. It is the most common single cause.
  • The method itself is the last suspect, not the first.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch scikit-learn

Verified with torch 2.5.1 and scikit-learn 1.7.2 on CPU. Downloads nothing.

One word, nineteen points

Here is a detail that looks like nothing and is not. Two optimisers, the same nominal setting, the same seed, the same data.

same_config_different_result.py
import torch
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split

d = load_digits()                       # 1797 tiny 8x8 images, ships with sklearn
X = torch.tensor(d.data / 16.0, dtype=torch.float32)
y = torch.tensor(d.target, dtype=torch.long)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=0)

def train(opt_name, wd, seed=0):
    torch.manual_seed(seed)
    net = torch.nn.Sequential(torch.nn.Linear(64, 64), torch.nn.ReLU(),
                              torch.nn.Linear(64, 10))
    make = torch.optim.Adam if opt_name == "Adam" else torch.optim.AdamW
    opt = make(net.parameters(), lr=1e-3, weight_decay=wd)
    lossf = torch.nn.CrossEntropyLoss()
    for _ in range(300):
        opt.zero_grad()
        lossf(net(Xtr), ytr).backward()
        opt.step()
    with torch.no_grad():
        acc = (net(Xte).argmax(1) == yte).float().mean().item()
        return acc, net[0].weight.norm().item()

for wd in (0.0, 0.1):
    for name in ("Adam", "AdamW"):
        acc, norm = train(name, wd)
        print(f"weight_decay={wd:<4} {name:6s} test accuracy {acc:.4f}   |W1| {norm:.2f}")
Output
weight_decay=0.0  Adam   test accuracy 0.9537   |W1| 11.42
weight_decay=0.0  AdamW  test accuracy 0.9537   |W1| 11.42
weight_decay=0.1  Adam   test accuracy 0.7611   |W1| 2.64
weight_decay=0.1  AdamW  test accuracy 0.9537   |W1| 11.28

With decay switched off, the two are byte-identical — proof that nothing else differs between the runs. Switch decay on with the same number, and one loses nineteen points.

Now imagine a paper whose config file says weight_decay: 0.1 and whose text says "we use Adam". You have no way to know which of these two behaviours produced their number, and the answer is worth nineteen points.

Why they diverge

Adam(weight_decay=w) adds w * theta to the gradient. That sum then travels through Adam's division by the square-rooted second moment, so the effective amount of decay applied to each parameter depends on the size of its gradient history. Parameters with small gradients get crushed.

AdamW(weight_decay=w) subtracts lr * w * theta from the parameter directly, after the Adam step, untouched by the division. That is the decoupling Loshchilov and Hutter proposed in 2019, and it is why the weight norm moves from 11.42 to 11.28 rather than to 2.64.

The two were the same word in most codebases for years. Papers written before the distinction was widely understood used whichever their framework provided.

The gap checklist, in order

Work down this list, changing one thing per run and recording the effect.

Evaluation. Best epoch or last epoch? Validation or test? Single run or averaged? Are you using the same metric implementation — macro or micro averaging, and which class treated as positive?

Data. Normalisation constants (per-dataset means, or 0.5?). The exact split, including whether the paper's "validation" set is your "test" set. Deduplication. Sorting order, which changes batch composition.

Schedule. Warmup length, decay shape, total steps. A cosine schedule cut short is a different schedule, not a shorter one.

Regularisation. Which parameters receive weight decay — most modern recipes exclude biases and normalisation parameters, and almost no paper says so. Dropout placement. Label smoothing. Gradient clipping threshold.

Reported quantity. Weight averaging or an exponential moving average of parameters is used routinely and mentioned rarely. It is typically worth several tenths of a point.

Numerics. Precision (float32, bfloat16, or mixed), and library version. Framework defaults change between releases; PyTorch's own default for TF32 matrix multiplication on recent GPUs changed in version 1.12, which moves results.

Common mistakes

Changing several things at once. The gap becomes uninterpretable the moment two variables move together. One change, one run, one line in a log.

Comparing your mean against their single number. Run five seeds and compare distributions. If their number sits inside your range, there is nothing left to explain — see seed variance and error bars.

Assuming the same argument name means the same thing. weight_decay is the example above. learning_rate differs by whether it is per-batch or per-sample. epochs differs by whether the count includes warmup.

Tuning your version until it matches. Search until the gap closes and you have fitted your hyperparameters to the paper's number, which is a form of test-set overfitting. Fix a protocol, then measure.

Giving up before writing down the residual. A documented gap with four explained components is a useful artefact. An abandoned attempt is not.

Try it yourself

Add a third configuration: AdamW with weight_decay=0.1 but the decay excluded from biases, by passing two parameter groups. Predict first whether the accuracy moves at all on a model this small, then check. Then set weight_decay=0.01 for plain Adam and find the value at which its behaviour becomes indistinguishable from AdamW's.

What to learn next

Researcher — Mathematics and papers.

Decompose the gap, do not close it

The productive framing is attribution, not repair. For each candidate detail $d$ with plausible values, measure the effect of switching it while holding everything else fixed, and record $\Delta_d$. The reproduction succeeds when $\sum_d \Delta_d$ accounts for the observed gap within seed noise, and it succeeds informatively when the list is published.

This is an ablation study run over configuration choices rather than model components, and it has the same requirement: change one factor per run. Where interactions are suspected — schedule length and weight decay interact strongly — a small factorial design over two or three factors is worth the compute.

State the residual with an error bar. "Remaining gap 0.4 points, seed standard deviation 0.5 points over five runs" is a complete result. "We could not reproduce it" is not.

Details known to move results by more than a point

Ordered roughly by observed impact in published reproduction reports:

  • Evaluation protocol. Test-time augmentation is the largest single lever in vision. Papers reporting multi-crop or multi-scale evaluation are not comparable to a single centre crop, and the difference is routinely more than a point. Check which the paper's table reports before anything else.
  • Weight-decay coupling and scope. The developer section shows the coupling. Scope — excluding biases and normalisation parameters — is nearly universal in practice and nearly absent from papers.
  • Learning-rate schedule and warmup. Total-step-dependent schedules make "same learning rate, fewer epochs" a substantively different experiment.
  • Batch size, coupled to learning rate. Linear scaling with warmup (Goyal et al., 2017) is the standard adjustment; applying a paper's learning rate at a different batch size without adjustment is a different optimiser configuration.
  • Parameter averaging. EMA of weights, or averaging the last $k$ checkpoints, appears in released code far more often than in text.
  • Preprocessing constants. Dataset-specific channel means and standard deviations versus generic 0.5 values.
  • Numerical precision and kernels. Mixed precision, TF32, cuDNN algorithm selection and non-deterministic reduction order all perturb results. Individually small; collectively enough to hide a real effect.

Reading a framework's formula rather than its argument name

The epsilon in Adam is a compact case study. The paper gives Algorithm 1, and then, immediately before Section 2.1, an algebraically rearranged form that is more efficient and places epsilon differently — outside the bias-corrected second moment rather than inside it. The two forms are not identical when epsilon is not negligible.

Frameworks implemented different forms. Keras documents its epsilon argument as the "epsilon hat" of that rearranged version, while PyTorch's implementation follows Algorithm 1 and adds epsilon to the bias-corrected square root. With the default of $10^{-8}$ the difference is immaterial; with the $10^{-3}$ to $10^{-6}$ values common in transformer and reinforcement-learning recipes, it is not.

The general rule: when a hyperparameter matters, read the reference implementation's arithmetic rather than its documentation, and confirm that your framework computes the same expression. Reading a research codebase is the skill that makes this cheap.

When the number was optimistic

Some gaps do not close because the target was drawn from the upper tail. Reported single-run numbers are frequently the best of several attempts, whether or not the paper says so, and benchmark test sets that a community has optimised against for years carry their own inflation (Recht et al., 2019).

Distinguish the cases empirically. If your distribution over seeds contains their number, the method reproduces and the report was optimistic. If your distribution sits entirely below theirs and every detail has been ablated, you have evidence of an unstated ingredient — the subject of the next lesson — or of a result that does not hold. Both are worth publishing, and the second requires the first to have been done exhaustively.

What to learn next