Too small a model, or trained too little?
Both problems show a training loss that is too high, and one column tells them apart — whether the training loss is still falling — so the fix takes minutes instead of a week of guessing.
- 13 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A training loss that is too high has two very different causes, and whether the line is still moving tells you which.
Packing for a month-long trip with a small cabin bag is one kind of problem. You can spend all evening rearranging, and the clothes will not fit. The bag is the limit.
Packing a huge suitcase in thirty seconds is a completely different problem. Give it another ten minutes and everything goes in comfortably.
Both end with clothes left on the bed. Only one is fixed by more time. The mistake that costs weeks is treating the small-bag problem as a not-enough-time problem, and buying more time again and again.
Why this had to be invented
"My model is not accurate enough" is the most common sentence in machine learning, and it points at four or five different repairs.
Two of them look identical on a chart. Capacity — how much detail a model is able to hold — being too small looks like a high loss. Training time being too short also looks like a high loss.
Guessing wrong is expensive in both directions. Training a too-small model for ten times longer burns a week and changes nothing. Buying a bigger model when the small one was still improving wastes money and makes everything slower forever.
The good news is that the two have completely different signatures, and the signature is free to read.
How it works
Look at one thing: is the training loss still falling?
TRAINED TOO LITTLE MODEL TOO SMALL
train loss train loss
\ \
\ \____________________
\ still heading down flat, and it has been
\ flat for a long time
fix: train longer, or raise fix: a bigger model, or
the learning rate better featuresIf the line is still descending, you stopped early. Keep going.
If the line has been flat for a long stretch and it is flat at a high value, more time will not help. The model has hit its ceiling.
There is a five-minute test that settles it beyond doubt. Take about fifty training examples and try to make the model memorise them perfectly. Any model with enough capacity can memorise fifty examples. If yours cannot get near-perfect on fifty, the problem is not time.
A real example you have seen
Autocorrect on an old phone versus a new one. The old one offers the same three wrong words no matter how long you use it. Small dictionary, simple rule, and that is its ceiling.
A modern keyboard learns your names, your slang and your two languages over a few weeks. Same task, different capacity, and only one of them benefits from more time.
Remember this
- Both problems show a high training loss. The chart alone is not enough.
- Still falling means train longer. Flat and high means the model is too small.
- The settling test: a model that cannot memorise fifty examples has a capacity or code problem, never a patience problem.
What to learn next
- Finding a learning rate — the cause that impersonates a capacity ceiling most often.
- Learning curves — the same question asked against data size rather than training time.
- Deciding a model is good enough — what to do once neither cause applies any more.
Developer — Code and libraries.
Setup
pip install torchVerified with torch 2.5.1 (CPU), Python 3.10. The first script trains nine models and takes about a minute on a laptop CPU. Nothing is downloaded.
Nine runs: three sizes by three budgets
The task has two interaction terms, so a very small network genuinely cannot represent it.
import torch, torch.nn as nn
def make_data(n, seed):
g = torch.Generator().manual_seed(seed)
X = torch.randn(n, 8, generator=g)
y = (X[:,0]*X[:,1] + 0.6*X[:,2]*X[:,3] - 0.5*X[:,4] > 0).float()
return X, y[:, None]
Xtr, ytr = make_data(2000, 0); Xva, yva = make_data(4000, 999)
def train(hidden, epochs, seed=0):
torch.manual_seed(seed)
net = nn.Sequential(nn.Linear(8, hidden), nn.ReLU(),
nn.Linear(hidden, hidden), nn.ReLU(), nn.Linear(hidden, 1))
opt, lf = torch.optim.Adam(net.parameters(), lr=0.01), nn.BCEWithLogitsLoss()
tr_hist = []
for _ in range(epochs):
opt.zero_grad(); l = lf(net(Xtr), ytr); l.backward(); opt.step(); tr_hist.append(l.item())
with torch.no_grad():
return tr_hist, lf(net(Xtr), ytr).item(), lf(net(Xva), yva).item()
print("hidden epochs train loss val loss drop over last 20% of epochs")
for h in [2, 8, 64]:
for ep in [100, 400, 1600]:
hist, tr, va = train(h, ep)
k = ep // 5
drop = hist[-k] - hist[-1] # the column that decides the diagnosis
print(f"{h:6d} {ep:6d} {tr:10.4f} {va:8.4f} {drop:27.4f}")hidden epochs train loss val loss drop over last 20% of epochs
2 100 0.6039 0.6143 0.0304
2 400 0.5018 0.5332 0.0031
2 1600 0.4970 0.5308 0.0001
8 100 0.0945 0.1257 0.0228
8 400 0.0405 0.1380 0.0050
8 1600 0.0220 0.2990 0.0021
64 100 0.0024 0.1023 0.0022
64 400 0.0001 0.1552 0.0001
64 1600 0.0000 0.2066 0.0000Read the last column first. It is the diagnosis, and the other columns are the story.
The walkthrough
hidden=2 is capacity-bound, and the numbers prove it. Sixteen times more training moves the training loss from 0.5018 to 0.4970. The last-20% drop collapses to 0.0001 — the line is flat. This model will never solve the task, and no patience, learning rate or scheduler changes that. Two hidden units cannot represent two interaction terms.
hidden=8 at 100 epochs is time-bound. Training loss is 0.0945 and still falling by 0.0228 over the last twenty epochs. That is a run stopped mid-descent, and the fix is more epochs, not more parameters.
The same hidden=8 at 1600 epochs shows the third problem. Training loss keeps improving (0.0220) while validation loss climbs to 0.2990. That is memorising, and it is the shape from how to read a loss curve. Notice that "train longer" was the correct advice at 100 epochs and the wrong advice at 1600 — the diagnosis has to be re-made, not made once.
The best model on this page is hidden=64 at 100 epochs, with validation loss 0.1023. The largest model and the shortest budget. If you had swept only one variable at a time from a bad starting point, you would probably have missed it. Capacity and training time interact, so sweep them together — a small grid beats a long line.
Watch out for a trap in the last column. hidden=64 at 100 epochs also has a small drop (0.0022), because its training loss has already reached 0.0024 and there is nothing left to fall. A flat curve at a low value means converged; a flat curve at a high value means capacity-bound. The height matters as much as the slope.
The five-minute settling test
When the table is ambiguous, ask a sharper question: can this architecture memorise fifty rows? Append this to the same file, so Xtr, ytr and the imports are already in scope.
Xs, ys = Xtr[:50], ytr[:50]
print("hidden train loss on 50 rows after 500 epochs verdict")
for h in [2, 8, 64]:
torch.manual_seed(0)
net = nn.Sequential(nn.Linear(8, h), nn.ReLU(),
nn.Linear(h, h), nn.ReLU(), nn.Linear(h, 1))
opt, lf = torch.optim.Adam(net.parameters(), lr=0.01), nn.BCEWithLogitsLoss()
for _ in range(500):
opt.zero_grad(); lf(net(Xs), ys).backward(); opt.step()
with torch.no_grad(): v = lf(net(Xs), ys).item()
print(f"{h:6d} {v:39.4f} {'cannot memorise 50 rows' if v > 0.05 else 'capacity is fine'}")hidden train loss on 50 rows after 500 epochs verdict
2 0.3423 cannot memorise 50 rows
8 0.0002 capacity is fine
64 0.0000 capacity is fineFifty rows is small enough that any adequate architecture should reach near-zero loss. Failing this test means one of three things: too little capacity, a learning rate that cannot make progress, or a bug in the data or loss wiring. All three are worth knowing about before you spend a night training. The full version of this technique is overfit one batch.
A decision list you can actually follow
- Run the fifty-row test. Fails? Fix capacity, learning rate or the bug before anything else.
- Passes, but the full training loss is high and still falling? Train longer, or use a learning-rate schedule.
- Passes, and the full training loss is high and flat? Capacity is enough for fifty rows and not enough for the real task. Increase width or depth, or improve the features.
- Training loss is low and validation loss is rising? Different problem entirely — that is memorising, and the fix is early stopping, regularisation or more data.
- Everything is low and you still are not happy? You are at the task's noise floor. That conversation belongs to deciding a model is good enough.
Common mistakes
Reading the validation curve to diagnose underfitting. Validation loss mixes capacity, training time and overfitting into one number. The training loss is the clean signal for "can this model represent the pattern at all".
Calling a plateau a ceiling too early. Some plateaus break — see plateaus, sudden drops and double descent. The fifty-row test tells them apart in minutes, because a capacity ceiling is visible on fifty rows and a plateau usually is not.
Adding parameters when the learning rate was the problem. A too-small learning rate produces a flat, high training loss that looks exactly like a capacity ceiling. Run the learning-rate finder before you buy a bigger model.
Sweeping one variable at a time. The winner on this page was the largest model with the shortest budget. A one-at-a-time search starting from hidden=8, epochs=400 would have walked away from it.
Assuming "capacity" only means parameter count. Depth, activation choice, normalisation and input features all change what a model can represent. Two hidden units fail here because of the interaction terms; adding X0*X1 as an input feature would fix the same model without adding one parameter.
Try it yourself
Add X[:,0]*X[:,1] and X[:,2]*X[:,3] as two extra input columns, so the network gets nine features instead of eight. Rerun the grid. Predict first: which cells improve most, and does hidden=2 become usable? The answer is the strongest argument for feature engineering you will find in this section.
What to learn next
- Finding a learning rate — the cause that impersonates a capacity ceiling most often.
- Learning curves — the same question asked against data size rather than training time.
- Deciding a model is good enough — what to do once neither cause applies any more.
Researcher — Mathematics and papers.
Separating approximation, optimisation and estimation error
The excess risk of a trained model decomposes into three terms, and this lesson is a procedure for attributing an observed gap to the right one. With $\mathcal{H}$ the hypothesis class, $h^\star$ the best possible predictor, $h^\star_{\mathcal{H}}$ the best member of $\mathcal{H}$, $\hat{h}$ the empirical risk minimiser and $\tilde{h}$ what the optimiser actually returned after a finite budget,
$$ R(\tilde h) - R(h^\star) = \underbrace{R(h^\star_{\mathcal{H}}) - R(h^\star)}{\text{approximation}} + \underbrace{R(\hat h) - R(h^\star{\mathcal{H}})}{\text{estimation}} + \underbrace{R(\tilde h) - R(\hat h)}{\text{optimisation}} . $$
"Model too small" is approximation error; "trained too little" is optimisation error; "memorising" is estimation error. The training loss isolates the first two from the third, because it contains no held-out sample, which is why the diagnostic column of the table is a training-loss slope rather than anything measured on validation data. Bottou and Bousquet (2008), The Tradeoffs of Large Scale Learning (NeurIPS), formalise the budgeted version of this decomposition and derive the counter-intuitive practical result that under a fixed compute budget the optimal choice is often a worse optimiser run on more data — the same reason the largest model with the shortest schedule wins on this page.
Why two hidden units genuinely cannot
Universal approximation (Cybenko, 1989; Hornik, 1991) guarantees arbitrary accuracy only as width grows without bound; it says nothing at width 2. For a ReLU network the function class is piecewise linear, and the number of linear regions is bounded above by a quantity that grows polynomially in width and exponentially in depth — Montúfar et al. (2014), On the Number of Linear Regions of Deep Neural Networks (NeurIPS). A target with an $x_0x_1$ interaction requires a decision boundary that a two-region-per-layer network cannot bend into shape, so the approximation term is bounded away from zero for reasons that have nothing to do with the optimiser.
This also explains the feature-engineering exercise: supplying $x_0x_1$ explicitly changes $\mathcal{H}$, collapsing the approximation term to near zero at unchanged parameter count. Capacity is a property of the model and its input representation, which is why "add parameters" and "add features" are alternative treatments for one diagnosis.
Compute allocation between size and steps
The trade-off this lesson explores by hand is the central question of large-scale training. Kaplan et al. (2020), Scaling Laws for Neural Language Models, fit loss as a joint power law in parameters, data and compute and concluded that under a fixed compute budget one should train very large models for relatively few steps and stop well short of convergence — a direct formalisation of "hidden=64, epochs=100". Hoffmann et al. (2022) (Chinchilla) revised the exponents, finding parameters and tokens should scale roughly in proportion, and demonstrated a 70B model trained on 1.4T tokens beating a 280B model trained on far fewer. The methodological lesson beneath both is the one in the walkthrough: the optimum lives on a two-dimensional surface, and a one-dimensional sweep can miss it entirely.
Practically, this argues for an IsoFLOP design — hold a compute budget fixed, vary the size/steps split along it, and repeat at two or three budgets — rather than a full grid. The envelope of those curves is what extrapolates; individual points do not.
The overfit-a-tiny-subset test, formally
Training to near-zero loss on $n = 50$ examples asks whether $\mathcal{H}$ can shatter a small sample under the actual optimiser. Zhang et al. (2017), Understanding Deep Learning Requires Rethinking Generalization (ICLR), established the empirical fact that standard architectures can fit random labels on full-size datasets, which makes failure on fifty real examples strongly diagnostic: it implies a defect in capacity, conditioning or wiring rather than an intrinsic difficulty of the task. The test is cheap, has near-zero false-positive rate in practice, and is the reason it belongs before every long run.
Its limitation is worth stating. Passing rules out gross capacity failure but says nothing about the sample complexity of the real task, so it can leave a genuine underfit undetected when the target requires many more effective parameters than fifty points can reveal. The complementary check is the shape of the training loss against training-set size — the learning curve — where a training loss that stays high as $n$ grows is the population-level signature of the same ceiling.
What to learn next
- Finding a learning rate — the cause that impersonates a capacity ceiling most often.
- Learning curves — the same question asked against data size rather than training time.
- Deciding a model is good enough — what to do once neither cause applies any more.