Validation and Hyperparameter Search

Learning curves

A learning curve plots model score against training-set size, and its shape answers the most expensive question in ML — will more data help?

On this page 5
  1. Why this had to be invented
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A learning curve shows how a model's score changes as you give it more training data.

Its shape tells you whether collecting more data is worth the money.

Before paying for more tuition classes, a sensible student checks one thing: plot mock-test marks against weeks of study. If the line has been flat since March, more of the same studying will not move it. If it is still climbing, keep going.

Models deserve the same check. "Get more data" is the most common advice in ML — and often the most expensive to follow. Labelling more scans, running more surveys, buying more storage. The learning curve is the graph that says whether the bill is worth paying, before you pay it.

Why this had to be invented

Two curves are drawn together: the score on the training data itself, and the score on held-out validation data, each measured at several training sizes. The space between them is where the diagnosis lives.

A big gap — training score high, validation score well below — means the model memorises its sample rather than learning the pattern: the overfitting signature. More data genuinely helps here, because memorising gets harder as the sample grows.

Both curves low and flat, hugging each other, is the opposite disease. The model is too crude for the pattern, and feeding a crude model more examples produces a well-fed crude model. Money spent on data here is wasted — the fix is a richer model or better features.

How it works

   overfitting model                underfitting model

score                            score
 1.0 ─ train ─────────────        1.0
        \ gap                            train ~~~~~~~~
         \                        0.7  ~~~~~~~~~~~~~~~ validation
          validation, rising            (flat, low, together)
          toward the train line

 more data → helps                more data → will NOT help

Read the validation curve's trend for the forecast, and the gap for the diagnosis.

A real example you have seen

Any team deciding whether to label 100,000 more images has faced this graph. Data labelling costs real money. The learning curve is the business case for or against it. The modern giant-scale version made headlines: AI labs plot exactly these curves to decide whether the next model deserves ten times more data.

Remember this

  • Learning curve = score versus training size, for training and validation data together.
  • Gap + rising validation → more data will help. Flat and low together → it will not.
  • Check the curve before paying for data, not after.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2 on CPU. Runs in about ten seconds.

Two models, two diagnoses, one dataset

learning_curves_demo.py
import numpy as np
from sklearn.datasets import make_classification
from sklearn.ensemble import RandomForestClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import learning_curve

X, y = make_classification(n_samples=2000, n_features=20, n_informative=10,
                           flip_y=0.1, random_state=10)

for name, model in [("random forest", RandomForestClassifier(random_state=10)),
                    ("logistic regression", LogisticRegression(max_iter=2000))]:
    sizes, train, val = learning_curve(model, X, y, cv=5,
                                       train_sizes=np.linspace(0.1, 1.0, 5))
    print(name)
    for n, tr, va in zip(sizes, train.mean(axis=1), val.mean(axis=1)):
        print(f"  {n:5d} rows   train {tr:.3f}   validation {va:.3f}   gap {tr - va:.3f}")
Output
random forest
    160 rows   train 1.000   validation 0.814   gap 0.186
    520 rows   train 1.000   validation 0.854   gap 0.146
    880 rows   train 1.000   validation 0.879   gap 0.121
   1240 rows   train 1.000   validation 0.892   gap 0.108
   1600 rows   train 1.000   validation 0.890   gap 0.110
logistic regression
    160 rows   train 0.819   validation 0.765   gap 0.054
    520 rows   train 0.821   validation 0.829   gap -0.008
    880 rows   train 0.821   validation 0.830   gap -0.008
   1240 rows   train 0.839   validation 0.833   gap 0.006
   1600 rows   train 0.842   validation 0.833   gap 0.009

Two textbook cases from one dataset. The forest memorises everything at every size (train 1.000 throughout) while its validation score climbs 0.814 → 0.890 — data was helping, though the last step's flattening says returns are shrinking. The logistic model plateaus at 0.83 by row 520 with no gap at all: feeding it more rows is provably pointless, and only a better model or richer features can move it.

The walkthrough

learning_curve handles the bookkeeping. For each of the five sizes it trains on a subset and scores via 5-fold cross-validation — 25 fits per model here. The returned train and val arrays are (sizes × folds); averaging over folds smooths split luck.

Train accuracy of 1.000 is information, not an error. An unconstrained forest can memorise any sample — the training curve pins to the ceiling and the entire diagnosis moves to the validation curve and the gap.

A negative gap (row 520 of the logistic model) means validation scored above training — pure fold noise. Read gaps within ±0.01 as zero rather than inventing stories about them.

Extrapolate with care. The forest's 0.892 → 0.890 step suggests a plateau near 0.89, but two points are thin evidence. When the decision is expensive, add sizes — log-spaced sizes (np.geomspace) reveal curve shape more efficiently than linear ones.

Common mistakes

Confusing learning curves with training curves. This lesson's curves vary training-set size. The loss-versus-epochs plots watched during neural network training are a different tool answering a different question ("has training converged?"), despite sharing a name.

One split per size. Score-versus-size measured on a single split inherits all the dice-roll variance of that split, at every point. Always cross-validate each size — learning_curve does by default.

Declaring a plateau from linear spacing. Improvements often continue at a slowing rate — each doubling of data buys a similar sliver. Linear size steps make that look flat. Check with log-spaced sizes before cancelling the data budget.

Diagnosing with a tuned-to-death model. Curves measured on settings tuned for the full dataset can mislead at small sizes (the regularisation that is right for 160 rows differs from 1,600). For big decisions, lightly re-tune per size.

Try it yourself

Give the forest max_depth=5 and rerun. Predict first: what happens to the training curve, the gap, and the validation plateau? Then swap in train_sizes=np.geomspace(0.05, 1.0, 8) and check whether the forest's "plateau" at 1,600 rows survives log-spaced scrutiny.

What to learn next

Researcher — Mathematics and papers.

Two curves, two limits

Fix a learning algorithm and let $E_{val}(n)$ and $E_{train}(n)$ denote expected validation and training error at training-set size $n$. Standard behaviour: $E_{val}(n)$ decreases toward an asymptote $E_\infty$; $E_{train}(n)$ increases toward the same asymptote (small samples are easier to fit than the population). The asymptote decomposes as irreducible noise plus the model class's approximation error — the bias term — while the gap $E_{val}(n) - E_{train}(n)$ tracks the variance/capacity term. The two diagnoses of the beginner section are readings of which term dominates.

Power-law shape

Empirically and under several theoretical models, the excess error follows

$$ E_{val}(n) - E_\infty \approx b\, n^{-\beta} $$

with $\beta$ typically in $[0.5, 1]$ for classical models under smoothness assumptions (Amari, Fujita and Shinomoto, 1992, give $\beta = 1$ regimes for parametric learning). Large-scale empirical confirmation: Hestness et al. (2017), Deep learning scaling is predictable, empirically — across vision, language and speech, held-out loss follows power laws over orders of magnitude of $n$, with domain-dependent exponents. The programme culminates in neural scaling laws — Kaplan et al. (2020), Scaling laws for neural language models, and the compute-optimal reallocation of Hoffmann et al. (2022, Chinchilla) — where fitted curves in $n$ (tokens) and parameters drive multi-million-dollar data-versus-model budget decisions. The same logic, at learn-more-Python scale, is this lesson's plateau reading; the connection to LLM practice is covered in how LLMs work.

Fitting and extrapolating curves

Practical sample-size forecasting fits $\hat{E}(n) = a + b n^{-\beta}$ to measured points by nonlinear least squares (weighted toward larger $n$, where variance is lower) and extrapolates. Figueroa et al. (2012, BMC Medical Informatics) validate this workflow for clinical text classification. Caveats: power-law fits are locally excellent and globally fallible — regime changes (new classes appearing, distribution shift as data collection broadens) break the extrapolation; report prediction intervals, not point forecasts. Mohr and van Rijn (2022), Learning curves for decision making in supervised machine learning — a survey, catalogue parametric forms and their failure modes.

Uses beyond the data-budget question

  • Early discarding in model selection: extrapolated curves let searches kill configurations predicted to plateau low (Domhan, Springenberg and Hutter, 2015) — the model-based cousin of successive halving's empirical ranking.
  • Data valuation: the marginal value of a data point at current $n$ is the local slope $\partial E / \partial n$ — the quantity that pricing data acquisition or active learning implicitly estimates.
  • Sanity-checking pipelines: a validation curve that fails to improve with $n$ on a task known to be learnable is a strong leakage or preprocessing-bug signal — sometimes the cheapest integration test a pipeline gets.

Sample-complexity theory, briefly

PAC bounds give worst-case guarantees of the form $n = O!\left(\frac{d + \log(1/\delta)}{\varepsilon^2}\right)$ samples for error $\varepsilon$ at confidence $1-\delta$, with $d$ the VC dimension. These are distribution-free and correspondingly loose — often orders of magnitude above empirical needs. The practical stance: theory bounds the shape of what is possible; measured learning curves estimate the constants that matter for your dataset. When they conflict, the curve wins.

What to learn next

What to learn next

These follow on from what you just read.

  • Validation and Hyperparameter Search

    AIC and BIC

    AIC and BIC score a model as fit minus a fee for complexity, so you can compare candidates from a single fit each — no refitting, no folds, and a different answer from each of the two.

  • Calibration and Uncertainty

    ROC vs precision-recall curves

    ROC curves judge a ranking against all the negatives while precision-recall curves judge what the alarms contain — on rare positives the two tell very different stories.

  • Calibration and Uncertainty

    Choosing a threshold from costs

    The 0.5 cut-off is a default nobody chose on purpose — when the two mistakes have different prices, the cheapest threshold follows from those prices, and it is rarely 0.5.