Reading Training Curves

When your validation curve is too noisy to trust

A small validation set makes a curve that wobbles by more than the improvement you are chasing — so before reading any dip, measure how wide the noise band is and refuse to act inside it.

On this page 5
  1. Why this had to be invented
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A small validation set makes the curve jump around by more than the improvements you are chasing.

Stand on a bathroom scale five times in a row. The number moves by a few hundred grams each time, and your body has not changed at all. The scale is telling the truth; it has a resolution, and the wobble is smaller than that resolution.

Your validation curve has a resolution too. It is set by how many examples you test on, and by the random choices made during training. Below that resolution, the curve is reporting weather, not climate.

Chasing a dip inside the noise band is how weeks disappear.

Why this had to be invented

The validation set is the pile of examples the model never learns from, used to check its progress. People keep it small because every example put in the validation pile is an example not available for training.

But a small pile means a shaky measurement. If you test on 100 examples and get 85 right, your model's true accuracy could honestly be anywhere from about 78% to about 92%. That range is wider than most improvements anyone chases.

Worse, the training itself is random. The starting weights are random. The order examples arrive in is random. Run the identical recipe twice and you get two different curves.

So before believing any wiggle, you have to know how big a wiggle means nothing.

How it works

  measured accuracy on 100 validation rows

  0.78 |=============== the true accuracy is somewhere
  0.85 |======#======== in this whole band
  0.92 |===============
              ^
        one measurement lands here

  A model that improves from 0.85 to 0.87 moves LESS
  than the width of the band. You cannot see it.

Two separate causes of wobble, and they need different fixes.

Measurement wobble comes from the validation set being small. Fix: use more validation examples, or check the same model on several different splits.

Training wobble comes from randomness inside training. Fix: run the same recipe with several different random starts, and read the group instead of one line.

A real example you have seen

Election exit polls. One channel says 42%, another says 46%, and both surveyed a few thousand people out of a hundred million voters.

Neither channel is lying. The honest ones print "margin of error plus or minus 3 points" underneath. Viewers then know that a 2-point gap between two parties is not a story. Your validation curve deserves the same footnote, and nobody ever writes it.

Remember this

  • A curve has a noise band. Measure it before reading anything inside it.
  • Small validation set means a wide band. 100 rows gives roughly 7 accuracy points of slack either way.
  • Run the same recipe with 3 to 5 different random seeds and read all the lines together.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch numpy

Verified with torch 2.5.1 (CPU), numpy 1.26.4, Python 3.10. The second script trains five models and takes about forty seconds on a laptop CPU. No downloads.

How wide is the band? Measure it directly

Fix a model whose real accuracy is exactly 0.85, then ask what a validation set of a given size would report.

noise_band.py
import numpy as np

rng = np.random.default_rng(0)
TRUE = 0.85                                    # the model's real accuracy

print(" val rows   spread of measured accuracy over 2000 resamples")
for n in [100, 500, 2000, 10000]:
    measured = rng.binomial(n, TRUE, size=2000) / n
    lo, hi = np.percentile(measured, [2.5, 97.5])
    print(f"{n:9d}   {lo:.3f} to {hi:.3f}   (band width {hi-lo:.3f})")
Output
 val rows   spread of measured accuracy over 2000 resamples
      100   0.780 to 0.920   (band width 0.140)
      500   0.820 to 0.880   (band width 0.060)
     2000   0.834 to 0.866   (band width 0.032)
    10000   0.843 to 0.857   (band width 0.014)

The model is identical in every row. Only the size of the ruler changes. With 100 validation examples, two models 5 accuracy points apart are statistically indistinguishable — and 100-example validation sets are extremely common in tutorials and small projects.

Note the shape of the improvement: going from 100 rows to 10,000 rows, a hundredfold increase, narrows the band tenfold. Halving your error bar costs four times the validation data.

Now the training wobble

Same recipe, five different seeds, nothing else changed.

seed_spread.py
import numpy as np, torch, torch.nn as nn

def make_data(n, seed):
    g = torch.Generator().manual_seed(seed)
    X = torch.randn(n, 8, generator=g)
    y = (X[:, 0]*X[:, 1] + 0.4*X[:, 2] + 0.3*torch.randn(n, generator=g) > 0).float()
    return X, y[:, None]

Xtr, ytr = make_data(1000, 0)
Xva, yva = make_data(200, 999)                  # a deliberately small validation set

def train(seed, epochs=60, bs=32):
    torch.manual_seed(seed)                     # changes init AND batch order
    net = nn.Sequential(nn.Linear(8, 32), nn.ReLU(), nn.Linear(32, 1))
    opt, loss_fn = torch.optim.Adam(net.parameters(), lr=0.01), nn.BCEWithLogitsLoss()
    hist = []
    for _ in range(epochs):
        perm = torch.randperm(len(Xtr))
        for i in range(0, len(Xtr), bs):
            idx = perm[i:i+bs]
            opt.zero_grad(); loss_fn(net(Xtr[idx]), ytr[idx]).backward(); opt.step()
        with torch.no_grad():
            hist.append(loss_fn(net(Xva), yva).item())
    return np.array(hist)

curves = np.stack([train(s) for s in range(5)])   # five identical recipes

print("epoch   seed 0    mean of 5 seeds    spread across seeds")
for e in range(0, 60, 6):
    col = curves[:, e]
    print(f"{e:5d}   {curves[0, e]:.4f}   {col.mean():15.4f}    {col.min():.4f} - {col.max():.4f}")

best = curves.argmin(axis=1)
print(f"\n'best epoch' chosen by each seed: {list(best)}")
print(f"best val loss per seed:          {[round(v, 4) for v in curves.min(axis=1)]}")
Output
epoch   seed 0    mean of 5 seeds    spread across seeds
    0   0.5856            0.5779    0.5688 - 0.5856
    6   0.3231            0.3186    0.3023 - 0.3320
   12   0.3243            0.3085    0.2971 - 0.3243
   18   0.3102            0.3193    0.3102 - 0.3274
   24   0.3179            0.3245    0.3179 - 0.3309
   30   0.3408            0.3365    0.3140 - 0.3486
   36   0.3528            0.3448    0.3269 - 0.3653
   42   0.3422            0.3467    0.3253 - 0.3637
   48   0.3684            0.3430    0.3218 - 0.3684
   54   0.3940            0.3536    0.3258 - 0.3940
Output
'best epoch' chosen by each seed: [5, 12, 7, 12, 12]
best val loss per seed:          [0.3018, 0.3016, 0.2984, 0.2988, 0.2971]

Exact numbers vary slightly across PyTorch builds and CPU vendors, because floating-point reduction order differs. The pattern will not vary, and the pattern is the point.

The walkthrough

The five best losses span 0.2971 to 0.3018 — a range of 0.0047. Any experiment claiming a 0.004 improvement from a clever change has measured nothing. This is the single most useful number on the page: it is the smallest difference your setup can detect, and you can compute it in forty seconds.

"Best epoch" ranges from 5 to 12. Same code, same data, same hyperparameters. If your write-up says "the model converges at epoch 5", it is a statement about one random seed, not about the model.

Seed 0's line is not monotone. It goes 0.3102 at epoch 18, up to 0.3408 at epoch 30, back to 0.3422 at 42, up to 0.3940 at 54. Reading that single line, epoch 42 looks like a recovery. The five-seed mean shows a flat drift upward with no recovery at all.

The spread column is your error bar, for free. Print min - max across seeds beside every curve, forever. It costs one extra column and it kills more bad conclusions than any other habit in this section.

Shrinking the noise, in order of value

  1. Use a bigger validation set before anything else. It is the only fix that improves the measurement itself rather than averaging over a bad one. If data is scarce, k-fold cross-validation reuses every row as validation exactly once.
  2. Run 3 to 5 seeds and report mean plus spread — see seed variance and error bars for the reporting conventions.
  3. Evaluate less often, on more data. Twenty honest measurements beat two hundred shaky ones, and validation is not free compute.
  4. Smooth only for display, never for decisions. The reasons are in how smoothing and log scales mislead you.

Common mistakes

Picking the single best epoch from a single run. That epoch's loss is the minimum of a noisy sequence, so it is biased low by construction — the more often you evaluate, the luckier the minimum looks. Report the loss at the selected epoch on a separate test set, or accept that the number is optimistic.

Comparing two runs that differ in a hyperparameter and in the seed. Then you cannot tell which one moved the curve. Hold the seed fixed when comparing, and separately measure seed spread to know what counts as a real difference.

Evaluating on a validation set with an unlucky class balance. With 200 rows and a 5% positive class you have 10 positives, and recall moves in steps of 10 percentage points. Stratify the split — cross-validation strategies covers the mechanics.

Blaming the noise when the pipeline is unstable. Dropout, data augmentation and batch-norm statistics all change evaluation results if the model is left in training mode. Confirm model.eval() is called (train and eval mode) before concluding the wobble is inherent.

Try it yourself

Change make_data(200, 999) to make_data(4000, 999) and rerun. Predict first: does the spread across seeds shrink, and by how much? Whatever remains after that change is training wobble, not measurement wobble — and it is the part a bigger validation set can never fix.

What to learn next

Researcher — Mathematics and papers.

Two variance components, and only one of them is fixable by data

Write the measured validation score of a run as

$$ \hat{M} = M(\theta(\omega)) + \varepsilon(V) $$

where $\omega$ collects the training randomness (initialisation, batch ordering, dropout masks, augmentation draws, non-deterministic kernel reductions), $\theta(\omega)$ is the resulting parameter vector, $M$ is the population metric, and $\varepsilon(V)$ is the sampling error of evaluating on the finite set $V$. The total variance decomposes as

$$ \operatorname{Var}(\hat{M}) = \underbrace{\operatorname{Var}\omega!\big(M(\theta(\omega))\big)}{\text{training variance}} + \underbrace{\mathbb{E}\omega\big[\operatorname{Var}(\varepsilon \mid \theta)\big]}{\text{evaluation variance}} . $$

Enlarging $V$ attacks only the second term, at rate $1/|V|$. For accuracy, $\varepsilon$ is binomial, so the standard error is $\sqrt{p(1-p)/|V|}$ — the source of the "±7 points at $n=100$" figure, since $1.96\sqrt{0.85 \cdot 0.15/100} \approx 0.070$. Averaging over seeds attacks only the first term, at rate $1/k$ for $k$ seeds. Neither operation substitutes for the other, and a paper reporting five seeds on a 200-row validation set has controlled the smaller of its two variances.

Bouthillier et al. (2021), Accounting for Variance in Machine Learning Benchmarks (MLSys), measure both components empirically across standard benchmarks and reach a counter-intuitive practical conclusion: randomising all sources of variation across a modest number of runs estimates the performance distribution better than fixing seeds and running more of the same configuration, and many published gains lie inside the resulting distribution. Reimers and Gurevych (2017), Reporting Score Distributions Makes a Difference (EMNLP), showed the same for sequence labelling — score differences reported as improvements were frequently within seed spread.

The selection bias in "best epoch"

Choosing $\hat{t} = \arg\min_t \hat{R}_V(\theta_t)$ and then reporting $\hat{R}V(\theta{\hat{t}})$ reports the minimum of a correlated noisy sequence, which is downward-biased. For $T$ evaluations with per-evaluation noise standard deviation $\sigma$ and weak correlation, the expected optimism grows roughly like $\sigma\sqrt{2\ln T}$ — increasing evaluation frequency inflates the reported score without improving the model. This is the same winner's curse that governs test-set overfitting and hyperparameter search, and it is why an unbiased estimate requires a set never used for selection.

Two mitigations with different costs. A three-way split (train / validation for selection / test for reporting) is clean and expensive in data. Nested cross-validation is data-efficient and expensive in compute; Varma and Simon (2006) (BMC Bioinformatics) quantify the bias of the non-nested alternative on small samples, where it is large enough to invert conclusions.

Deciding whether a difference is real

The correct test is paired: evaluate both models on the same validation examples and analyse the per-example differences, which removes the shared difficulty of the examples from the variance. McNemar's test is the standard choice for paired binary correctness; Dietterich (1998), Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms (Neural Computation), evaluates five candidate tests for Type I error and recommends McNemar or 5×2-fold cross-validated $t$ over the resampled $t$-test, which is badly miscalibrated because its folds are not independent. The bootstrap over validation examples gives a confidence interval on the difference without distributional assumptions and generalises to metrics with no closed-form variance (bootstrapping).

For curves specifically, the object of inference is a whole trajectory, and pointwise intervals under-cover when read across many epochs. Reporting a band from independent seeds and refusing to interpret crossings inside it is the cheap, defensible practice; statistical significance in ML and comparing models fairly carry the full apparatus.

Non-determinism you did not choose

Even with every seed fixed, exact reproduction can fail. Atomic accumulation order on GPU, cuDNN algorithm selection, TF32 and mixed-precision rounding, and multi-worker data loading all introduce run-to-run differences (random seeds and reproducibility). Forcing deterministic kernels costs throughput and does not remove the scientific problem — a result that survives only under one bit-exact configuration was never robust. Determinism is a debugging tool; seed spread is the scientific report.

What to learn next