Learning rate
In one sentence The learning rate is the size of each training step — the single hyperparameter most likely to make or break a run.
Updated
The learning rate is how big a step the optimizer takes each time it updates the weights.
You are walking to a doorway in a dark room. Huge strides get you across the room fast — and past the doorway, again and again, forever. Tiny shuffles guarantee you arrive, sometime next week. The learning rate is your stride length, and both extremes fail in their own recognisable way.
The failure signatures are worth memorising, because you diagnose them from the loss curve alone:
too high : loss jumps around, or climbs, or becomes NaN
slightly high : loss falls fast, then flattens above where it should
good : steady fall, gradually levelling out
too low : loss creeps down painfully slowlyTypical starting points: around 1e-3 for Adam from scratch, and 1e-5 to 1e-4 when fine-tuning a pretrained model — small, because big steps would trample what the model already knows.
In serious training the rate is not constant. A schedule changes it over time: a brief warmup from near zero, then a slow decay, often cosine-shaped. Warmup exists because the earliest steps, taken with randomly initialised statistics, are the easiest place to wreck a run.
Where to go next
- Full lesson: Gradient descent
- Related terms: gradient-descent, adam, hyperparameter, loss-function