Reading and Reimplementing Papers

Reading a research codebase without drowning

Never read a research repo front to back — arrive with one question, find the entry point, the config and the loss, and let grep do the rest.

On this page 7
  1. Arrive with a question
  2. The four things worth finding
  3. How it works
  4. The move that beats reading
  5. The honest bit
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Read a research codebase with one specific question in hand, and read only the parts that answer it.

Walk into a large railway station for the first time and you do not study the building. You look for one board — departures — find your platform number, and walk there. The station has fifty other things in it. None of them are your problem today.

People open a research repository and start at the top of the file list. Twenty minutes later they are inside a logging utility, three levels deep, with no idea how they got there.

Arrive with a question

Before opening anything, write down what you want to know. One sentence.

"What learning rate schedule did they use?" "Where is the loss computed?" "How is the data normalised?"

The learning rate is the size of each training step. A schedule is how that size changes over the run.

A question turns an unbounded reading task into a search. Searching is fast. Reading is slow. The whole skill is converting one into the other.

The four things worth finding

Almost every research repository has the same four load-bearing parts, whatever the folder names.

  1. The entry point. The file you run. Usually train.py, main.py, or something in scripts/.
  2. The config. Where the numbers live — learning rates, sizes, epochs. Often a .yaml or .json file, or a block of command-line defaults.
  3. The loss. Where the method's idea becomes arithmetic. This is the paper.
  4. The data path. What is loaded, and what is done to it before the model sees it.

Find those four and you understand the repository well enough for almost any purpose.

How it works

   your question
        │
        ▼
   README  →  entry point  →  config  →  the number you wanted
                   │
                   └─→  the loss function  →  the method itself

   everything else — logging, checkpointing, distributed
   training, argument plumbing — is NOT your problem

That last line is the hard part. Research repositories contain a great deal of machinery unrelated to the idea. Walking past it is a skill, not laziness.

The move that beats reading

Run it small.

Set the batch size to 2, the dataset to a handful of examples, the training length to one step. Then print things. Ten minutes of running the code teaches more than two hours of reading it. Running shows you what the code does. Reading shows you what it appears to do.

The honest bit

Research code is often messy, and that is not a character flaw. It was written under a deadline by people whose job was the idea, not the software. Dead branches, unused options and confusing names are normal.

Expecting it to be clean will make you doubt yourself. It is not you.

Remember this

  • Arrive with one question, and search rather than read.
  • Find the entry point, the config, the loss, the data path. Ignore the rest.
  • Run it tiny and print things. Behaviour beats appearance.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Verified with torch 2.5.1 on CPU. The line numbers below come from that version and will move in others — the technique is the transferable part, not the numbers.

A real codebase you already have

You do not need to clone anything to practise this. PyTorch's Adam implementation is on your disk, it is real production research-lineage code, and it answers a question from earlier in this section: exactly where does epsilon enter the update?

navigate_a_repo.py
import inspect
import re
import torch
import torch.optim.adam as adam_module

src = inspect.getsource(adam_module)
lines = src.splitlines()
print(f"torch {torch.__version__}")
print(f"{inspect.getsourcefile(adam_module).split('site-packages')[-1]}: {len(lines)} lines")

print("\ntop-level functions -- the map of the file:")
for name in re.findall(r"^def (\w+)", src, re.M):
    print("   ", name)

print("\nwhere does epsilon enter the update?")
for i, line in enumerate(lines, 1):
    s = line.strip()
    if s.startswith("#") or "eps" not in s:
        continue
    if "denom" in s or "add_(eps)" in s:
        print(f"  line {i}: {s}")
Output
torch 2.5.1+cu121
\torch\optim\adam.py: 803 lines

top-level functions -- the map of the file:
    _single_tensor_adam
    _multi_tensor_adam
    _fused_adam
    adam

where does epsilon enter the update?
  line 298: eps (float, optional): term added to the denominator to improve
  line 428: denom = (max_exp_avg_sqs[i].sqrt() / bias_correction2_sqrt).add_(eps)
  line 430: denom = (exp_avg_sq.sqrt() / bias_correction2_sqrt).add_(eps)

Eight hundred lines, reduced to the two that matter, in under a second.

The function list is the map. Three implementations of the same algorithm — one plain, one batched across tensors, one fused into a single kernel — plus a dispatcher named adam that chooses between them. That structure is extremely common in research code: a readable reference implementation beside optimised variants. Read the readable one and ignore the rest.

Line 298 is documentation, caught by the search and discarded on sight. Lines 428 and 430 are the answer. Line 428 sits in the AMSGrad branch, using the running maximum of the second moment; line 430 is the standard path.

The answer to the question. exp_avg_sq.sqrt() / bias_correction2_sqrt is the square root of the bias-corrected second moment, and .add_(eps) puts epsilon outside it. That is Algorithm 1 of Kingma and Ba, confirming what the previous lesson claimed about where PyTorch places it. The question is now settled by evidence rather than by documentation.

The same moves at the shell

For a repository you have cloned, the equivalents are one-liners. No output block follows this one on purpose — every line prints whatever that repository happens to contain, so the result is different for every reader.

bash
ls                                  # README, setup, and the folder names
cat README.md | head -60            # how do they say to run it?
find . -name "*.yaml" -o -name "*.json" | grep -i conf   # where the numbers live
grep -rn "def forward" --include=*.py .     # the model
grep -rn "loss" --include=*.py . | head     # the method
python train.py --help              # every knob, with its default
git log --oneline | head -20        # what changed recently

python train.py --help is the most underrated of these. Argument parsers document defaults that no README mentions, and those defaults are frequently the hyperparameters used for the paper.

Reading git history for the "why"

The code tells you what. Version history tells you why.

git log -S "weight_decay" finds every commit that added or removed that string. git blame path/to/file.py shows who last touched each line and in which commit, and the commit message often explains a choice that the code cannot. When a released repository has a tag or release matching the paper, check out that tag — the current main branch may be two years of unrelated evolution ahead of what produced the published numbers.

Common mistakes

Reading top to bottom. The file order is alphabetical, not logical. Enter through the entry point and follow calls.

Trusting names. A file called utils.py frequently contains the method. A function called normalize may not normalise. Verify by running, not by reading the identifier.

Assuming the repo matches the paper. Released code often reflects a later, different version of the work. Where they disagree, note the disagreement and prefer the code for reproducing numbers — then say which you followed.

Reading without running. Print shapes at the point where data enters the model. Two printed shapes settle questions that a day of reading leaves open.

Following every import. Depth-first reading is how people drown. When a call goes somewhere uninteresting, note the name and come back only if the question demands it.

Try it yourself

Answer a new question with the same technique: does PyTorch's SGD apply momentum before or after weight decay? Point the script at torch.optim.sgd, search for momentum and weight_decay in the update lines, and write the update as a formula. Then check your formula against the docstring at the top of the same file.

What to learn next

Researcher — Mathematics and papers.

Reading for provenance, not comprehension

For reproduction work the question is rarely "how does this code work?" It is "which code produced Table 3?" Those are different questions, and the second is answered from history rather than from source.

  • Pin to the artefact. Prefer a release tag, a commit hash cited in the paper, or an archived snapshot with a DOI. A repository's default branch is a moving target, and a reproduction that does not name a commit is not reproducible.
  • Diff the paper's era against HEAD. git log --since around the submission date, and git diff <tag>..HEAD -- <method file>, expose whether the method itself changed after publication.
  • Look for the config that matches the reported run. Many repositories carry a configs/ directory in which one file corresponds to each table row. That file settles more questions than the paper's method section.
  • Check the issue tracker. Open issues titled "cannot reproduce Table 2" are the fastest available summary of where a codebase and its paper disagree, and often contain a maintainer's answer.

Known classes of paper–code divergence

Divergences between released code and published description recur in predictable forms, and each has a standard place to look:

  • Extra tricks in the code. Gradient clipping, EMA of weights, label smoothing, warmup — present in the training loop, absent from the text. Grep the optimiser step and the training loop for anything that touches parameters outside the described update.
  • Different evaluation. Test-time augmentation, a different metric implementation, or model selection on the test split. Read the evaluation function before the model.
  • Data differences. Filtering, deduplication, or a different split than the standard one for that dataset. Read the dataset class end to end; it is usually short and disproportionately important.
  • Hyperparameters that differ from the paper's table. Common, and usually resolved in favour of the code.

Documenting these divergences is a legitimate research output. Several of the field-scale audits discussed in reading a results table sceptically began as exactly this kind of careful code reading.

Instrumenting rather than reading

For a large codebase, three techniques beat further reading:

  • Hooks over inspection. Register forward and backward hooks to record the shapes and statistics of every intermediate tensor in one pass. See forward and backward hooks.
  • Trace the call graph on a one-step run. python -m trace --listfuncs or a profiler on a single batch shows which functions actually execute, which is typically a small fraction of the repository.
  • Assert your model of the code. Insert assertions expressing what you believe — that a tensor is normalised, that a mask is boolean, that no gradient flows through a branch. A failed assertion is a corrected belief, obtained in seconds.

What to write down

Keep a reading log with one line per finding: the question, the file and line, the commit hash, and the answer. It becomes the provenance section of your reproduction report, it survives the repository changing under you, and it prevents the most expensive failure mode in this work — answering the same question twice, three weeks apart, and getting different answers without noticing.

What to learn next