Early stopping
In one sentence Early stopping ends training when performance on held-out data stops improving, keeping the model from sliding into memorisation.
Updated
Early stopping watches the model's score on data it does not train on, and stops training when that score stops improving.
It is the kitchen rule for boiling milk: the useful phase and the disaster phase look similar until you cross the line, so you watch the pot and act at the turn. Training loss will keep falling almost forever. The number that matters is the loss on a held-out validation-set, and at some point it turns upward. That turn is the model shifting from learning to memorising.
The mechanism is plain bookkeeping. After each epoch, evaluate on validation data. Keep a copy of the best weights seen so far. If validation loss has not improved for a set number of epochs — called the patience — stop, and restore that best copy.
epoch: 1 3 5 7 9 11
train : 0.90 0.60 0.43 0.30 0.21 0.14
valid : 0.93 0.65 0.51 0.48 0.52 0.60
▲
stop here, keep these weightsEarly stopping is free regularization: no formula changes, no retraining, one saved checkpoint. The one caution is patience set too low, which stops training during a temporary plateau that the model would have pushed through.
Where to go next
- Full lesson: Overfitting and underfitting
- Related terms: overfitting, epoch, validation-set, regularization