Reading and Reimplementing Papers
Reading an ML paper in three passes
Reading a paper front to back is the slowest way to learn nothing — three passes of increasing depth let you reject most papers in ten minutes and understand the rest properly.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Read a research paper in three passes: a ten-minute skim to decide if it is worth your time, an hour to understand what it did, and a long session to rebuild it.
Think about how you handle a new recipe. First you glance at it — how long, what ingredients, is this a weeknight dish or a Sunday project. Most recipes get closed right there. Only the ones that survive the glance get read properly.
Nobody reads a cookbook cover to cover. Yet almost everyone opens their first research paper at word one and grinds forward until they give up somewhere in section 3.
Why the front-to-back habit fails
A paper is not written to be read in order. It is written to survive review.
The introduction repeats the abstract. The related-work section exists to reassure reviewers that nobody was ignored. The proofs sit in the middle because a journal wanted them there. Meanwhile the one thing you need is scattered. What changed sits in an algorithm box on page 3. Whether it helped sits in a table on page 7.
Reading in order means burning your best attention on the least useful pages. By the time you reach the part that matters, you are tired and no longer sure why you started.
The three-pass method comes from a short 2007 note by S. Keshav, written for computer-science students drowning in exactly this problem. It works because each pass has one job and a stopping rule.
How it works
Pass 1 (10 min) title, abstract, headings, figures, conclusion
→ QUESTION: is this worth another hour?
most papers stop here. that is the point.
Pass 2 (1 hour) read the body properly, skip the proofs
→ QUESTION: could I explain this to a friend?
Pass 3 (a day) rebuild the idea yourself, challenge every choice
→ QUESTION: would I get the same result?The first pass is a filter, not a summary. You are allowed — encouraged — to abandon papers. Reading five papers badly teaches less than reading one properly and dropping four fast.
What to pull out on the first pass
Four things, and nothing else:
- The claim. One sentence: what does this do better than what came before?
- The evidence. What was measured, on what data, against what alternative?
- The cost. More compute, more memory, more code, more tuning?
- The reusable bit. Is there an algorithm box or an equation you could copy?
If you cannot find the claim in ten minutes, that is information about the paper.
A real example you have seen
Every popular AI method you use daily arrived as a paper someone had to read this way. The optimiser that trains almost every model today came from a nine-page paper in 2015. Thousands of engineers read its algorithm box, copied it, and shipped it. Most never opened the convergence proof. That proof turned out to have a flaw nobody noticed for three years.
Remember this
- Three passes: filter, understand, rebuild. Each has its own stopping rule.
- Pass 1 exists so you can abandon most papers quickly and guiltlessly.
- Find the claim, the evidence, the cost, the reusable bit — before reading anything else.
What to learn next
- Reading a results table sceptically — the pass-1 skill that saves the most time.
- Decoding the notation — for when pass 2 stalls on symbols.
- A staged plan for reimplementing a paper — what pass 3 looks like when done for real.
Developer — Code and libraries.
The worked example for this whole section
We will use one real paper throughout: Kingma and Ba (2015), "Adam: A Method for Stochastic Optimization" (arXiv:1412.6980, ICLR 2015). It is a good teaching paper because it is short, its core is a single algorithm box, and you already use it every time you write torch.optim.Adam.
Here is a first pass on it.
- The claim. An optimiser combining per-parameter step sizes with momentum, needing little memory, suited to noisy or sparse gradients and non-stationary objectives.
- The evidence. Section 6: logistic regression on MNIST and on IMDB bag-of-words features, multilayer networks on MNIST, a convolutional network on CIFAR-10, plus a study of the bias-correction term on a variational autoencoder. Training cost curves against SGD with momentum, AdaGrad, RMSProp and AdaDelta.
- The cost. Two extra arrays the size of the parameters, one for each running average. Four hyperparameters, with defaults given.
- The reusable bit. Algorithm 1, ten lines of pseudocode, plus defaults: step size 0.001, decay rates 0.9 and 0.999, epsilon 1e-8.
Ten minutes of work, and you now know whether to spend the hour.
Setup
pip install torchVerified with torch 2.5.1 on CPU. No download, no GPU, runs in under a second.
Testing a sentence from the abstract
The abstract says the method is invariant to diagonal rescaling of the gradients. That is a concrete, checkable claim, and checking one claim from the abstract is the cheapest possible third pass.
The test: minimise the same problem twice, once with all coordinates on the same scale and once with them stretched wildly. If the claim holds, the path taken should barely change.
import torch
def run(opt_name, scales, steps=60):
"""Minimise f(w) = sum_i scale_i * w_i^2, starting from w = 1."""
w = torch.full((3,), 1.0, requires_grad=True)
s = torch.tensor(scales)
opt = (torch.optim.Adam([w], lr=0.1) if opt_name == "adam"
else torch.optim.SGD([w], lr=0.1))
for _ in range(steps):
opt.zero_grad()
(s * w**2).sum().backward()
opt.step()
return w.detach().numpy().round(3)
for name in ("adam", "sgd"):
print(f"{name} scales 1, 1, 1 -> w = {run(name, [1.0, 1.0, 1.0])}")
print(f"{name} scales 1, 100, 0.01 -> w = {run(name, [1.0, 100.0, 0.01])}")adam scales 1, 1, 1 -> w = [-0.035 -0.035 -0.035] adam scales 1, 100, 0.01 -> w = [-0.035 -0.035 -0.035] sgd scales 1, 1, 1 -> w = [0. 0. 0.] sgd scales 1, 100, 0.01 -> w = [0. nan 0.887]
The claim survives. Adam's three coordinates land in an identical place whether the problem is balanced or stretched by a factor of ten thousand. Plain gradient descent, at the same step size, blows up on the steep coordinate and barely moves the flat one.
That took eleven lines. You now understand why the method exists, which no amount of re-reading the introduction would have given you.
The walkthrough
Every coordinate starts at 1 and the minimum is at 0, whatever the scales are. So the only thing that changes between the two runs is how the gradient is scaled per coordinate — exactly the transformation the abstract claims to be immune to.
nan is the honest output, not a bug. With a scale of 100 the gradient is 200 * w, and a step size of 0.1 overshoots the minimum by more than it started. The iterate doubles in size each step until it overflows. Reproducing a failure is as informative as reproducing a success.
Adam is invariant here, not exactly invariant in general. The epsilon term in the denominator breaks perfect invariance when gradients get very small. The paper's claim is about the mechanism; verify claims at the strength they were made.
Round the output. Printing full float precision invites you to over-read tiny differences that come from floating-point ordering, not from the method.
Common mistakes
Reading the related-work section first. It is a map of a field you do not know yet, written in names. Read it on pass 2 or 3, when the names have meaning.
Refusing to abandon a paper. Sunk cost applies to reading. If pass 1 says the claim is vague or the evidence is thin, close it. There is another paper.
Skipping the figures. Figures and their captions carry more information per second than any paragraph. On pass 1, read every caption before reading any body text.
Believing you understood without testing anything. The feeling of understanding a method and the ability to reproduce it are different states. Eleven lines of code separates them.
Try it yourself
Do a pass 1 on any paper you have been avoiding. Write the four bullets — claim, evidence, cost, reusable bit — in a text file, and set a ten-minute timer. Then find one sentence in its abstract that is concrete enough to test, and test it.
What to learn next
- Reading a results table sceptically — the pass-1 skill that saves the most time.
- Decoding the notation — for when pass 2 stalls on symbols.
- A staged plan for reimplementing a paper — what pass 3 looks like when done for real.
Researcher — Mathematics and papers.
What the third pass is actually for
Keshav's third pass is described as a virtual re-implementation: reconstruct the work from the same assumptions, then compare against what the authors did. The value is in the divergences. Every place your reconstruction differs from the paper is either a detail the paper omitted, or an assumption the authors made without saying so.
Three targets on that pass:
- The assumptions behind each theorem. Bounded gradients, convexity, a particular decay schedule. Papers often prove under conditions their own experiments violate.
- The gap between the theory and the experiments. Theory on convex problems, experiments on deep networks, is a very common pairing and a very weak one.
- What the authors chose not to compare against. Absent baselines are the loudest part of a results table, and the subject of the next lesson.
The worked example, continued
Adam's Theorem 4.1 gives an $O(\sqrt{T})$ regret bound for online convex optimisation under bounded gradients and a decaying step size $\alpha_t = \alpha / \sqrt{t}$ — none of which describe the deep-network experiments in Section 6, where a constant step size is used on non-convex objectives.
Three years later, Reddi, Kale and Kumar (2018), On the Convergence of Adam and Beyond (ICLR 2018 best paper), identified an error in that analysis and constructed a simple convex problem on which Adam converges to the wrong point. Their fix, AMSGrad, enforces a non-increasing effective step size by keeping a running maximum of the second-moment estimate.
Two lessons sit inside that history. First, a published proof is a claim like any other. Second, the flaw did not stop the method from being useful — Adam remained the default optimiser throughout, because the empirical evidence was strong and independent of the proof. Separating "the theory is wrong" from "the method does not work" is a research skill in itself.
Building the map around a paper
A paper read alone is nearly worthless; a paper read in its citation neighbourhood is a position in an argument. Practical moves:
- Forward citations. Search who cited it. Replications, corrections and "X considered harmful" papers appear here, and they are the fastest route to knowing whether a result held up.
- Reviews, where public. OpenReview carries the reviewer objections and author rebuttals for most ICLR, NeurIPS and ICML submissions since roughly 2018. The reviewers usually found the weakness before you did.
- Versions. arXiv keeps every version. Diffing v1 against v3 shows what review forced the authors to change, which is often the part they were least confident about.
- The code, if it exists. Covered in reading a research codebase — the resolution of most ambiguity in the text.
Reading for a purpose, not for completeness
Depth should be set by what you plan to do:
| Your goal | Deepest pass needed |
|---|---|
| Know the method exists | 1 |
| Explain it in a reading group | 2 |
| Use a library implementation | 2, plus the defaults table |
| Reimplement from scratch | 3 |
| Extend it or review it | 3, plus the proofs |
Most engineering work stops at pass 2. Deciding in advance which row you are in prevents both under-reading and the more common failure of reading a paper deeply for no reason.
What to learn next
- Reading a results table sceptically — the pass-1 skill that saves the most time.
- Decoding the notation — for when pass 2 stalls on symbols.
- A staged plan for reimplementing a paper — what pass 3 looks like when done for real.