Overfitting and underfitting
Overfitting is a model that memorised the training data instead of learning from it. Underfitting is a model too simple to learn the pattern at all.
- 15 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Overfitting is when a model memorises the practice questions instead of learning the subject.
Think of two students preparing for an exam. The first one finds last year's question paper and learns all the answers by heart. The second one barely opens the book and hopes for the best.
Hand them this year's paper — with new questions — and both fail. The memoriser is overfitting. The one who never studied is underfitting.
The student you want is the third one. They worked through the practice paper and understood why each answer was right. Now they can handle questions they have never seen.
Why this is the central problem
Here is the thing that makes machine learning hard.
You can always make a model score perfectly on the data you trained it on. Give it enough freedom and it will memorise every single row, mistakes and all. Your score will look magnificent.
Then you show it one new example and it falls apart.
The score you actually care about is on data the model has never seen. That score cannot be improved by trying harder. Almost every technique in machine learning exists to manage this one gap.
The two failures, side by side
| Underfitting | Overfitting | |
|---|---|---|
| The student | Never studied | Memorised the paper |
| On practice questions | Does badly | Does perfectly |
| On the real exam | Does badly | Does badly |
| The model is | Too simple | Too free |
| The fix | Give it more freedom | Rein it in, or get more data |
Notice the giveaway in the middle row. Both end up failing the real exam, but they look completely different during practice. That difference is how you tell them apart.
How it looks
error
^
|\ /
| \ /
| \ /
| \ test error /
| \ /
| \_____ ___/
| \______________/
|
| \
| \____
| \________ training error ________
+---------------------------------------------> model freedom
UNDERFIT GOOD OVERFIT
both errors both errors training error tiny,
are high are low test error climbingTraining error keeps falling as you give the model more freedom. It never turns around.
Test error falls, reaches a low point, and then starts climbing. That turning point is what you are hunting for.
This part confuses almost everyone the first time. Read it twice — that is completely normal.
Here is the key idea. A falling training error can mean real progress. It can also mean memorisation. The training error by itself cannot tell you which one is happening.
How you catch it
You hide some data from the model.
Before training, put aside part of your data and do not let the model see it. Train on the rest. Then test on the hidden part.
All your data
|
+---> 80% -> the model trains on this
|
+---> 20% -> locked away, used only to checkIf the model does well on both parts, it learned something. If it does well on the training part and badly on the hidden part, it memorised. This is the single most useful habit in the whole field.
Where you have seen it
- A face unlock that only works in your bedroom light. It overfit to the lighting it was set up in.
- A voice assistant that understands one accent. It overfit to the accents in its training data.
- A prediction app that was brilliant in testing and useless once launched. Very often overfitting, discovered too late.
- A spam filter that blocks nothing new. Too simple, so underfitting.
What to do about it
If your model is overfitting:
- Get more data. This is the most reliable fix and usually the hardest.
- Give the model less freedom. A simpler model has less room to memorise.
- Stop training earlier, before memorisation starts.
- Remove features that are not helping.
If your model is underfitting:
- Give it more freedom, or pick a more capable model.
- Add better features, so the pattern is easier to find.
- Train for longer.
Remember this
- Overfitting means memorising the training data. Underfitting means never learning the pattern.
- Always keep some data hidden from the model, and score on that.
- A perfect training score is a warning sign, never a celebration.
What to learn next
- Train, test and validation splits — how to hide data correctly.
- Model evaluation — the measurements that reveal these problems.
- Random forest — a model built specifically to resist overfitting.
Developer — Code and libraries.
Setup
pip install scikit-learn numpyWatching a model memorise
We generate points along a wave and add noise. Then we fit curves of increasing bendiness. Each one is scored twice — on data it trained on, and on data it never saw.
import numpy as np
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import PolynomialFeatures
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error
rng = np.random.default_rng(0)
x = np.linspace(0.0, 1.0, 20)
y = np.sin(2 * np.pi * x) + rng.normal(0, 0.25, size=20)
# Odd-numbered points are held back so the model never sees them while learning
x_train, y_train = x[0::2].reshape(-1, 1), y[0::2]
x_test, y_test = x[1::2].reshape(-1, 1), y[1::2]
for degree in (1, 3, 12):
model = make_pipeline(PolynomialFeatures(degree), LinearRegression())
model.fit(x_train, y_train)
train_err = mean_squared_error(y_train, model.predict(x_train)) ** 0.5
test_err = mean_squared_error(y_test, model.predict(x_test)) ** 0.5
print(f"degree {degree:2d} | train error {train_err:6.3f} | test error {test_err:8.3f}")degree 1 | train error 0.631 | test error 0.635 degree 3 | train error 0.194 | test error 0.230 degree 12 | train error 0.000 | test error 20.581
Reading the three rows
Degree 1 — underfitting. A straight line cannot follow a wave. Training error and test error are both high, and they are close to each other. That closeness is the signature of underfitting. The model is failing equally everywhere, which means the problem is the model, not the data split.
Degree 3 — about right. Both errors dropped. Test error (0.230) sits a little above training error (0.194), which is normal and healthy. A small honest gap is what a good model looks like.
Degree 12 — overfitting, spectacularly. Training error is 0.000. The model passes through all ten training points exactly. Test error is 20.581, which is roughly ninety times worse than the degree-3 model.
Sit with that contrast. Judged on training error alone, degree 12 is flawless and looks like the best model. Judged on data it has never seen, it is worse than useless. Both statements are about the same model, at the same moment.
Why degree 12 could hit zero
Ten training points and thirteen coefficients. There are more knobs than constraints, so an exact fit through every point is available. Between those points the curve is free to swing wildly, and it does. That is why predictions on the in-between test points explode.
This is the mechanical heart of overfitting. Enough parameters relative to data means memorisation becomes possible, and optimisers will happily find it.
Note on reproducibility
np.random.default_rng(0) fixes the noise, so these numbers reproduce exactly on any machine with a current NumPy. Swap in np.random.seed() with the legacy np.random.normal and you will get different values. If you change the seed, expect the degree-12 test error to move a great deal. It is a wild curve, so its error is unstable by nature. The pattern across the three rows will hold regardless.
Catching it properly with cross-validation
A single train/test split is noisy when data is small. Cross-validation splits the data several ways and averages the results.
import numpy as np
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import PolynomialFeatures
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import cross_val_score, KFold
rng = np.random.default_rng(0)
x = np.linspace(0.0, 1.0, 20).reshape(-1, 1)
y = np.sin(2 * np.pi * x.ravel()) + rng.normal(0, 0.25, size=20)
# shuffle=True matters here: x is sorted, and unshuffled folds would
# ask the model to extrapolate into a stretch of x it never saw
folds = KFold(n_splits=5, shuffle=True, random_state=0)
for degree in (1, 3, 5, 12):
model = make_pipeline(PolynomialFeatures(degree), LinearRegression())
# negative RMSE, because scikit-learn scorers are "higher is better"
scores = -cross_val_score(model, x, y, cv=folds,
scoring="neg_root_mean_squared_error")
print(f"degree {degree:2d} | mean CV error {scores.mean():8.3f}")degree 1 | mean CV error 0.745 degree 3 | mean CV error 0.228 degree 5 | mean CV error 0.227 degree 12 | mean CV error 28.454
Three things worth noticing.
Degree 1 is far too simple, and degree 12 is catastrophic. That much matches the single split.
Degrees 3 and 5 come out at 0.228 and 0.227. Treat those as tied. A gap in the third decimal place, across five folds of twenty points, is noise rather than evidence. Reporting "degree 5 is the best model" would be false precision. When two models are this close, pick the simpler one.
Most of these numbers are also higher than the single-split errors above. That is odd at first glance, because each fold trains on sixteen points rather than ten. More training data would normally mean lower error.
The reason is that the earlier split was unusually kind. Taking every other point puts each test point neatly between two training points, so every prediction is a short interpolation. Random folds sometimes remove two neighbouring points at once, leaving a real gap to guess across. Cross-validation is the harsher and more realistic estimate here. Compare models within one scheme, never across schemes.
The shuffle=True argument is doing real work. Passing cv=5 alone uses KFold without shuffling. Since x runs in sorted order, each fold would then hold out one contiguous stretch of the wave. The model would be scored on extrapolation rather than interpolation, and every error would inflate. On this data that pushes the degree-12 mean error into the thousands.
Common mistakes
Choosing a model by training score. Degree 12 has a perfect training score. This is the mistake the whole lesson exists to prevent.
Tuning on the test set. Try twenty models, keep the one with the best test score, and that score is now optimistic. You have leaked the test set into your decision. Use a three-way split, or cross-validate on training data and touch the test set once at the end.
Scaling before splitting. Calling StandardScaler().fit_transform(X) on all data before splitting leaks test statistics into training. Put the scaler inside a Pipeline, which scikit-learn refits correctly within each fold.
Assuming more data always fixes it. More data helps overfitting. It does nothing for underfitting — a straight line fitted to a million wave points is still a straight line.
Shuffling time-series data. train_test_split shuffles by default. On time-ordered data this lets the model train on the future and test on the past. Use TimeSeriesSplit.
Try it yourself
In overfit_demo.py, change the sample count from 20 to 200. It appears twice: in the linspace call, and in the size argument of the noise. Change nothing else — same model, same degrees.
degree 1 | train error 0.528 | test error 0.520 degree 3 | train error 0.262 | test error 0.240 degree 12 | train error 0.224 | test error 0.268
Degree 12 went from a test error of 20.581 to 0.268. It is now a reasonable model.
Nothing about the model changed. One hundred training points against thirteen coefficients leaves no room to memorise. The same flexible curve is now forced to find the actual pattern.
You have watched more data cure overfitting without touching the algorithm. This is why data collection is so often the highest-value work available to you.
Notice also that degree 1 barely improved (0.635 to 0.520). More data does nothing for underfitting, exactly as the common-mistakes section warned.
What to learn next
- Train, test and validation splits — splitting strategies and their failure modes.
- Model evaluation — choosing metrics that expose these problems.
- Random forest — averaging many overfit trees into one that generalises.
Researcher — Mathematics and papers.
The bias-variance decomposition
For squared loss, expected test error at a point x decomposes exactly into three terms.
E[ ( y - f_hat(x) )^2 ] = Bias[f_hat(x)]^2 + Var[f_hat(x)] + sigma^2where
Bias[f_hat(x)] = E[f_hat(x)] - f(x)
Var[f_hat(x)] = E[ ( f_hat(x) - E[f_hat(x)] )^2 ]f(x)— the true underlying functionf_hat(x)— the model fitted on one random training sampleE[.]— expectation over the draw of training setssigma^2— irreducible noise variance iny, unbeatable by any modelBias^2— systematic error from the hypothesis class being too restrictiveVar— sensitivity of the fitted model to which training sample it received
Underfitting is the high-bias regime. Overfitting is the high-variance regime. The sigma^2 term is the floor, and no algorithm goes below it.
Two important caveats. This decomposition is exact for squared loss. It does not carry over cleanly to 0-1 loss — see Domingos (2000) for the attempts. Also, "bias-variance tradeoff" describes a tendency, not a theorem. Modern methods can reduce both at once.
Capacity control
Classical theory bounds the generalisation gap by a capacity measure of H.
R(h) <= R_emp(h) + Omega(H, n, delta)Omega— a complexity penalty; VC-dimension bounds, Rademacher complexity, or covering numbersn— sample size, with the penalty typically shrinking asO(1/sqrt(n))
Structural Risk Minimisation (Vapnik) operationalises this: order nested hypothesis classes by capacity and pick the level minimising the bound. Regularisation is the continuous relaxation of the same idea.
Ridge / L2 weight decay: min L(theta) + lambda * ||theta||_2^2
Lasso / L1: min L(theta) + lambda * ||theta||_1
Early stopping: halt before empirical risk is minimised
Dropout: randomly zero units at training time
Data augmentation: enlarge the effective sample via invariancesTwo of these have exact equivalences worth knowing. L2 regularisation is the MAP estimate under a Gaussian prior on theta; L1 corresponds to a Laplace prior.
Early stopping has a similar dual. On a linear least-squares problem, stopping gradient descent early approximates ridge regression. The number of steps plays the role of 1/lambda.
Double descent
The classical U-shaped test-error curve is incomplete. This is the most important correction to the textbook picture in the last decade.
Belkin et al. (2019) documented double descent. The interpolation threshold is the capacity at which the model exactly fits the training data. As capacity grows past it, test error rises to a peak, then falls again. It often ends below the classical minimum.
test
error
^
| classical interpolation
| U-curve threshold
| /\ |
| / \ /|\
| / \____ / | \
|/ \____ / | \______________
| \________/ | modern regime
+--------------------------------------------------> capacity
p = nNakkiran et al. (2021) showed the effect appears across model size, training epochs, and dataset size in standard deep networks. The prevailing explanation involves implicit bias. Beyond interpolation there are infinitely many exact fits available. The optimiser's preference for minimum-norm solutions selects the smooth ones.
The practical consequence is genuine: for heavily over-parameterised models, "reduce capacity to stop overfitting" can be exactly the wrong move. The classical advice remains correct in the under-parameterised regime, which is where most tabular work still lives.
Benign overfitting
A related line of work asks how interpolating models can generalise at all. Bartlett et al. (2020) characterise benign overfitting in linear regression. When the covariance spectrum has many small eigenvalues, the minimum-norm interpolant absorbs noise into unimportant directions. It still predicts well.
This is not permission to ignore overfitting. The conditions are specific, and outside them interpolation is harmful in exactly the classical way.
Estimating the gap
| Method | Bias of the estimate | Cost | Notes |
|---|---|---|---|
| Single hold-out | Pessimistic; high variance | 1 fit | Fine for large n |
| k-fold CV | Slightly pessimistic | k fits | k = 5 or 10 standard |
| Leave-one-out | Nearly unbiased | n fits | High variance; closed form for linear models |
| Bootstrap .632+ | Corrected | ~200 fits | Efron & Tibshirani (1997) |
| Nested CV | Unbiased under tuning | k * m fits | Required when hyperparameters are tuned |
Nested cross-validation is under-used and matters. Tuning hyperparameters with the same cross-validation loop used to report performance produces an optimistically biased estimate. Cawley & Talbot (2010) documented how widespread this error is in the published literature.
Key references
- Geman, S., Bienenstock, E. & Doursat, R. (1992). Neural Networks and the Bias/Variance Dilemma. Neural Computation 4(1).
- Vapnik, V. (1998). Statistical Learning Theory. Wiley. Structural risk minimisation.
- Srivastava, N. et al. (2014). Dropout: A Simple Way to Prevent Neural Networks from Overfitting. JMLR 15.
- Zhang, C. et al. (2017). Understanding Deep Learning Requires Rethinking Generalization. ICLR.
- Belkin, M., Hsu, D., Ma, S. & Mandal, S. (2019). Reconciling Modern Machine Learning Practice and the Classical Bias-Variance Trade-off. PNAS 116(32).
- Nakkiran, P. et al. (2021). Deep Double Descent: Where Bigger Models and More Data Hurt. JSTAT.
- Bartlett, P., Long, P., Lugosi, G. & Tsigler, A. (2020). Benign Overfitting in Linear Regression. PNAS 117(48).
- Cawley, G. & Talbot, N. (2010). On Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation. JMLR 11.
What to learn next
- Model evaluation — estimator variance and metric selection.
- Train, test and validation splits — nested CV and leakage-safe protocols.
- Optimization — implicit regularisation of gradient descent.