ML Code That Survives

Getting out of the notebook

Notebooks are for exploring; scripts are for repeating — here is how to promote notebook code into functions, a main guard and arguments without losing what made the notebook useful.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A notebook is a kitchen experiment; a script is the written recipe — one is for discovering, the other is for repeating.

While inventing a new dish you taste constantly, adjust, backtrack, leave four half-used bowls on the counter. Wonderful — that mess is discovery. But you would never hand a stranger your messy counter and say "make this dish". You would write a recipe card: steps in order, quantities fixed, nothing left from earlier attempts.

A notebook — the cell-by-cell coding environment where you run pieces in any order — is the counter. A script — a file that runs top to bottom, the same way every time — is the recipe card. Both are good. Confusing them is the problem.

Why it exists

Notebooks reward exploration and quietly punish repetition. Cells run out of order, so the notebook's visible code and its actual state drift apart. A variable from a deleted cell lives on invisibly. "Restart and run all" fails on notebooks that "worked" for weeks.

Then the notebook becomes the team's training pipeline anyway, because it worked once. Every retraining turns into a ritual: open it, run cells in the remembered order, skip cell 7, hope. Rituals are not pipelines.

How it works

notebook (exploring)             script (repeating)

cell: load data                  def load_data(...):
cell: try stuff        ──→       def build_features(...):
cell: more tries                 def train(...):
cell: the good version           main() calls them, in order,
(run in who-knows what order)    top to bottom, every time

The promotion is a translation, not a rewrite. Keep the code that won, wrap the pieces in named functions, and fix an order. Throw away the failed attempts; they already served their purpose.

A real example you have seen

Scientists keep rough lab notebooks — scratches, crossed-out numbers, arrows. But the published method section is rewritten so any lab can repeat the procedure. No journal accepts a photo of the messy notebook. The rewrite is the science becoming shareable.

Remember this

  • Notebooks explore; scripts repeat. Use each for its job.
  • Out-of-order cells mean a notebook's code and its results can silently disagree.
  • Promote winning code into functions in a script; delete the failed attempts proudly.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Verified with scikit-learn 1.7.2, numpy 1.26.4, Python 3.10, CPU. Seeded, so your numbers should match.

The promoted script, whole

This is the target shape: the surviving notebook code as three functions plus an entry point with arguments.

train.py
import argparse
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

def load_data(seed):
    rng = np.random.default_rng(seed)
    X = rng.normal(size=(200, 4))
    y = (X[:, 0] + X[:, 1] + rng.normal(0, 0.8, 200) > 0).astype(int)
    return X, y

def train(X, y, seed):
    Xtr, Xte, ytr, yte = train_test_split(X, y, random_state=seed)
    model = LogisticRegression().fit(Xtr, ytr)
    return model, model.score(Xte, yte)

def main():
    parser = argparse.ArgumentParser()
    parser.add_argument("--seed", type=int, default=0)
    args = parser.parse_args()
    X, y = load_data(args.seed)
    model, acc = train(X, y, args.seed)
    print(f"seed {args.seed}: accuracy {acc:.3f}")

if __name__ == "__main__":
    main()

Run it twice, from the terminal:

bash
python train.py
python train.py --seed 7
Output
seed 0: accuracy 0.880
seed 7: accuracy 0.780

Two commands, two clean runs, a fresh process each time — no leftover state possible. And the ten-seed loop from seed variance is now a shell loop away.

The promotion procedure

  1. Restart and run all. Fix the notebook until it survives top-to-bottom execution. This alone flushes out the ghost variables.
  2. Name the phases. Almost every ML notebook is: load → features → train → evaluate. Wrap each phase's surviving code in a function with inputs and outputs — no globals crossing between them.
  3. Move functions to a file; leave calls behind. Paste the functions into train.py (inside src/, per the layout lesson). The notebook can now from churn.train import load_data — it becomes a user of the code instead of its container.
  4. Add the main() and the __main__ guard. The guard — if __name__ == "__main__": — means "run this only when executed directly, not when imported". It is what lets tests and notebooks import your functions without triggering a training run.
  5. Promote hardcoded values to arguments or a config. The seed above became --seed; real projects graduate to config files.

What the notebook remains for

Do not delete the notebook — demote it, to the job it is great at: plotting the trained model's errors, poking at data, trying the next idea by importing the script's functions. One direction of dependency (notebook imports script, never the reverse) keeps exploration fast and the pipeline clean.

Common mistakes

Promoting by exporting. "Download as .py" gives you the whole archaeology — every dead cell, in cell order. Promote by selecting the survivors, not by exporting the dig site.

Functions that secretly use globals. A function reading df from the surrounding notebook works until it is moved. Pass everything as parameters; the move then works mechanically.

One giant main() doing everything. That is the notebook again, wearing a function costume. The value is in the seams — load/features/train/evaluate as separately callable, separately testable units.

Waiting for the "right time" to promote. The right time is the first time anyone (including cron, including you-tomorrow) needs to run it again. After that, every ritual run adds risk.

Try it yourself

Take a real notebook of yours. Run "restart and run all" and count the failures — each is a ghost-state bug you were living with. Then promote it by the five steps and time yourself; the second promotion you ever do will be twice as fast.

What to learn next

Researcher — Mathematics and papers.

The measured state of notebooks

Pimentel et al. (2019), A Large-Scale Study on the Quality and Reproducibility of Jupyter Notebooks, MSR: of ~1.16M notebooks scraped from GitHub, roughly one in four executed end-to-end without error, and only about 4% reproduced their stored outputs. Dominant causes: missing dependencies, absent data files, and out-of-order execution — 36% of notebooks with numbered cells had them in non-linear order. The ghost-state problem is not folklore; it is the modal condition of public notebooks.

Head et al. (2019), Managing Messes in Computational Notebooks, CHI: studied the mess directly and built code-gathering tools — program slicing over the notebook's execution log to extract the minimal cell subset producing a chosen result. That is the "select the survivors" step, formalised: promotion is a slice of the notebook's dependency graph, not a copy of its text.

Grus (2018), I Don't Like Notebooks (JupyterCon) — the culture-war reference: hidden state, poor testability, and the anti-pattern of notebooks-as-production. The synthesis position this lesson takes — notebooks for exploration, scripts for repetition, one-way imports — is roughly where the community argument settled.

Why scripts are the reproducibility unit

A script run is a fresh interpreter: its result is a function of (code, inputs, environment, seed) — the provenance coordinates of reproducing your own result. A notebook's result is additionally a function of execution history, an unrecorded coordinate that can take factorially many values over $n$ cells. Restart-and-run-all is precisely the operation that collapses the history coordinate; scripts make that collapse permanent and mechanical.

The __main__ guard has a second, less-known justification: multiprocessing on Windows (and macOS spawn-mode) re-imports the main module in child processes; unguarded top-level training code forks recursively. PyTorch DataLoader workers hit exactly this — see DataLoader workers and speed.

Middle grounds and tooling

  • jupytext: two-way sync between .ipynb and plain .py with cell markers — notebooks become diffable and reviewable; promotion becomes a rename plus cleanup.
  • papermill: parameterise and execute notebooks headlessly — the "notebook as script" compromise; used in production at scale (Netflix's published pipeline architecture), at the cost of keeping the format's other liabilities.
  • nbdev (fast.ai): literate-programming discipline where the notebook is the source and export is automated, with tests in cells — a coherent alternative culture with its own tooling demands.
  • nbstripout / pre-commit hooks: strip outputs before commit — the minimum hygiene for notebooks in version control, since stored outputs both bloat diffs and (per Pimentel) rarely match re-execution anyway.

The selection pressure across all of these is identical: move code toward fresh-process, top-to-bottom, parameterised execution. Whichever tool gets you there, the property is the point.

What to learn next