Validation and Hyperparameter Search
Learning curves
A learning curve plots model score against training-set size, and its shape answers the most expensive question in ML — will more data help?
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A learning curve shows how a model's score changes as you give it more training data.
Its shape tells you whether collecting more data is worth the money.
Before paying for more tuition classes, a sensible student checks one thing: plot mock-test marks against weeks of study. If the line has been flat since March, more of the same studying will not move it. If it is still climbing, keep going.
Models deserve the same check. "Get more data" is the most common advice in ML — and often the most expensive to follow. Labelling more scans, running more surveys, buying more storage. The learning curve is the graph that says whether the bill is worth paying, before you pay it.
Why this had to be invented
Two curves are drawn together: the score on the training data itself, and the score on held-out validation data, each measured at several training sizes. The space between them is where the diagnosis lives.
A big gap — training score high, validation score well below — means the model memorises its sample rather than learning the pattern: the overfitting signature. More data genuinely helps here, because memorising gets harder as the sample grows.
Both curves low and flat, hugging each other, is the opposite disease. The model is too crude for the pattern, and feeding a crude model more examples produces a well-fed crude model. Money spent on data here is wasted — the fix is a richer model or better features.
How it works
overfitting model underfitting model
score score
1.0 ─ train ───────────── 1.0
\ gap train ~~~~~~~~
\ 0.7 ~~~~~~~~~~~~~~~ validation
validation, rising (flat, low, together)
toward the train line
more data → helps more data → will NOT helpRead the validation curve's trend for the forecast, and the gap for the diagnosis.
A real example you have seen
Any team deciding whether to label 100,000 more images has faced this graph. Data labelling costs real money. The learning curve is the business case for or against it. The modern giant-scale version made headlines: AI labs plot exactly these curves to decide whether the next model deserves ten times more data.
Remember this
- Learning curve = score versus training size, for training and validation data together.
- Gap + rising validation → more data will help. Flat and low together → it will not.
- Check the curve before paying for data, not after.
What to learn next
- Overfitting and underfitting — the two diseases this instrument diagnoses.
- Data labelling — what the "more data" branch actually costs.
- AIC and BIC — model selection from a single fit, when retraining many times is off the table.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7.2 on CPU. Runs in about ten seconds.
Two models, two diagnoses, one dataset
import numpy as np
from sklearn.datasets import make_classification
from sklearn.ensemble import RandomForestClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import learning_curve
X, y = make_classification(n_samples=2000, n_features=20, n_informative=10,
flip_y=0.1, random_state=10)
for name, model in [("random forest", RandomForestClassifier(random_state=10)),
("logistic regression", LogisticRegression(max_iter=2000))]:
sizes, train, val = learning_curve(model, X, y, cv=5,
train_sizes=np.linspace(0.1, 1.0, 5))
print(name)
for n, tr, va in zip(sizes, train.mean(axis=1), val.mean(axis=1)):
print(f" {n:5d} rows train {tr:.3f} validation {va:.3f} gap {tr - va:.3f}")random forest
160 rows train 1.000 validation 0.814 gap 0.186
520 rows train 1.000 validation 0.854 gap 0.146
880 rows train 1.000 validation 0.879 gap 0.121
1240 rows train 1.000 validation 0.892 gap 0.108
1600 rows train 1.000 validation 0.890 gap 0.110
logistic regression
160 rows train 0.819 validation 0.765 gap 0.054
520 rows train 0.821 validation 0.829 gap -0.008
880 rows train 0.821 validation 0.830 gap -0.008
1240 rows train 0.839 validation 0.833 gap 0.006
1600 rows train 0.842 validation 0.833 gap 0.009Two textbook cases from one dataset. The forest memorises everything at every size (train 1.000 throughout) while its validation score climbs 0.814 → 0.890 — data was helping, though the last step's flattening says returns are shrinking. The logistic model plateaus at 0.83 by row 520 with no gap at all: feeding it more rows is provably pointless, and only a better model or richer features can move it.
The walkthrough
learning_curve handles the bookkeeping. For each of the five sizes it trains on a subset and scores via 5-fold cross-validation — 25 fits per model here. The returned train and val arrays are (sizes × folds); averaging over folds smooths split luck.
Train accuracy of 1.000 is information, not an error. An unconstrained forest can memorise any sample — the training curve pins to the ceiling and the entire diagnosis moves to the validation curve and the gap.
A negative gap (row 520 of the logistic model) means validation scored above training — pure fold noise. Read gaps within ±0.01 as zero rather than inventing stories about them.
Extrapolate with care. The forest's 0.892 → 0.890 step suggests a plateau near 0.89, but two points are thin evidence. When the decision is expensive, add sizes — log-spaced sizes (np.geomspace) reveal curve shape more efficiently than linear ones.
Common mistakes
Confusing learning curves with training curves. This lesson's curves vary training-set size. The loss-versus-epochs plots watched during neural network training are a different tool answering a different question ("has training converged?"), despite sharing a name.
One split per size. Score-versus-size measured on a single split inherits all the dice-roll variance of that split, at every point. Always cross-validate each size — learning_curve does by default.
Declaring a plateau from linear spacing. Improvements often continue at a slowing rate — each doubling of data buys a similar sliver. Linear size steps make that look flat. Check with log-spaced sizes before cancelling the data budget.
Diagnosing with a tuned-to-death model. Curves measured on settings tuned for the full dataset can mislead at small sizes (the regularisation that is right for 160 rows differs from 1,600). For big decisions, lightly re-tune per size.
Try it yourself
Give the forest max_depth=5 and rerun. Predict first: what happens to the training curve, the gap, and the validation plateau? Then swap in train_sizes=np.geomspace(0.05, 1.0, 8) and check whether the forest's "plateau" at 1,600 rows survives log-spaced scrutiny.
What to learn next
- Overfitting and underfitting — the two diseases this instrument diagnoses.
- Data labelling — what the "more data" branch actually costs.
- AIC and BIC — model selection from a single fit, when retraining many times is off the table.
Researcher — Mathematics and papers.
Two curves, two limits
Fix a learning algorithm and let $E_{val}(n)$ and $E_{train}(n)$ denote expected validation and training error at training-set size $n$. Standard behaviour: $E_{val}(n)$ decreases toward an asymptote $E_\infty$; $E_{train}(n)$ increases toward the same asymptote (small samples are easier to fit than the population). The asymptote decomposes as irreducible noise plus the model class's approximation error — the bias term — while the gap $E_{val}(n) - E_{train}(n)$ tracks the variance/capacity term. The two diagnoses of the beginner section are readings of which term dominates.
Power-law shape
Empirically and under several theoretical models, the excess error follows
$$ E_{val}(n) - E_\infty \approx b\, n^{-\beta} $$
with $\beta$ typically in $[0.5, 1]$ for classical models under smoothness assumptions (Amari, Fujita and Shinomoto, 1992, give $\beta = 1$ regimes for parametric learning). Large-scale empirical confirmation: Hestness et al. (2017), Deep learning scaling is predictable, empirically — across vision, language and speech, held-out loss follows power laws over orders of magnitude of $n$, with domain-dependent exponents. The programme culminates in neural scaling laws — Kaplan et al. (2020), Scaling laws for neural language models, and the compute-optimal reallocation of Hoffmann et al. (2022, Chinchilla) — where fitted curves in $n$ (tokens) and parameters drive multi-million-dollar data-versus-model budget decisions. The same logic, at learn-more-Python scale, is this lesson's plateau reading; the connection to LLM practice is covered in how LLMs work.
Fitting and extrapolating curves
Practical sample-size forecasting fits $\hat{E}(n) = a + b n^{-\beta}$ to measured points by nonlinear least squares (weighted toward larger $n$, where variance is lower) and extrapolates. Figueroa et al. (2012, BMC Medical Informatics) validate this workflow for clinical text classification. Caveats: power-law fits are locally excellent and globally fallible — regime changes (new classes appearing, distribution shift as data collection broadens) break the extrapolation; report prediction intervals, not point forecasts. Mohr and van Rijn (2022), Learning curves for decision making in supervised machine learning — a survey, catalogue parametric forms and their failure modes.
Uses beyond the data-budget question
- Early discarding in model selection: extrapolated curves let searches kill configurations predicted to plateau low (Domhan, Springenberg and Hutter, 2015) — the model-based cousin of successive halving's empirical ranking.
- Data valuation: the marginal value of a data point at current $n$ is the local slope $\partial E / \partial n$ — the quantity that pricing data acquisition or active learning implicitly estimates.
- Sanity-checking pipelines: a validation curve that fails to improve with $n$ on a task known to be learnable is a strong leakage or preprocessing-bug signal — sometimes the cheapest integration test a pipeline gets.
Sample-complexity theory, briefly
PAC bounds give worst-case guarantees of the form $n = O!\left(\frac{d + \log(1/\delta)}{\varepsilon^2}\right)$ samples for error $\varepsilon$ at confidence $1-\delta$, with $d$ the VC dimension. These are distribution-free and correspondingly loose — often orders of magnitude above empirical needs. The practical stance: theory bounds the shape of what is possible; measured learning curves estimate the constants that matter for your dataset. When they conflict, the curve wins.
What to learn next
- Overfitting and underfitting — the two diseases this instrument diagnoses.
- Data labelling — what the "more data" branch actually costs.
- AIC and BIC — model selection from a single fit, when retraining many times is off the table.