Optimisers, Schedulers and the Training Loop
Finding a learning rate that works
The LR range test raises the learning rate step by step until training breaks, and the curve it draws tells you where a good value lives.
- 8 min read
- 3 reading levels
- Published
Read these first
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The fastest way to find a good learning rate is to raise it gradually until training breaks, then settle a comfortable distance below the breaking point.
It works like adjusting an unfamiliar shower tap. You turn the knob a little — nothing changes. A little more — pleasantly warm. Keep turning and suddenly it scalds.
Nobody finds the right setting by guessing one position and stepping in. You sweep the knob, feel where things go wrong, and settle safely below that point.
Why it exists
The learning rate is the single most influential setting in training. Too small and a model that could finish overnight takes weeks. Too large and the model never improves at all — it thrashes, or breaks completely.
The painful part: the good range differs between models, datasets and optimisers. There is no universal number. So instead of a number, you learn a procedure that finds the range in one cheap run.
How it works
loss
│████ ██
│ ████ ██
│ ██████ ██
│ ██████ ██
│ ████████████
└────────────────────────────────────── learning rate (growing)
too small: good zone: too big:
nothing moves loss falls loss explodesRun a short training where the learning rate grows every step. Watch the loss. It stays flat, then falls, then explodes. The good values sit in the falling zone, safely left of the explosion.
Where you have seen this
The button-press version ships in real tools: the fastai library made "find my learning rate" a one-line command, and PyTorch Lightning includes an automatic tuner. Both run exactly this sweep behind the scenes.
Remember this
- The learning rate is the most important knob in training.
- Sweep it upward in one cheap run; the loss curve shows the good zone.
- Pick a value in the steep-descent part, well below the explosion.
What to learn next
- Learning rate schedulers — turning your found value into a full schedule.
- When the loss will not go down — diagnosing the case where no learning rate seems to help.
- Tracking down a NaN loss — what the far side of the cliff looks like in a real run.
Developer — Code and libraries.
Setup
pip install torchOutputs captured with torch 2.5.1 on CPU; this run is seeded and deterministic.
The range test, hand-rolled
Twenty-two steps, doubling the learning rate each time:
import torch
import torch.nn as nn
torch.manual_seed(0)
X = torch.randn(512, 20)
true_w = torch.randn(20, 1)
y = X @ true_w + 0.1 * torch.randn(512, 1)
model = nn.Linear(20, 1)
opt = torch.optim.SGD(model.parameters(), lr=1e-5)
loss_fn = nn.MSELoss()
lr = 1e-5
for step in range(22):
for group in opt.param_groups:
group["lr"] = lr # raise the lr a little every step
opt.zero_grad()
loss = loss_fn(model(X), y)
loss.backward()
opt.step()
print(f"step {step:2d} lr={lr:10.6f} loss={loss.item():14.3f}")
lr *= 2.0 # double it and see what breaksstep 0 lr= 0.000010 loss= 8.909 step 1 lr= 0.000020 loss= 8.909 step 2 lr= 0.000040 loss= 8.908 step 3 lr= 0.000080 loss= 8.907 step 4 lr= 0.000160 loss= 8.904 step 5 lr= 0.000320 loss= 8.899 step 6 lr= 0.000640 loss= 8.888 step 7 lr= 0.001280 loss= 8.867 step 8 lr= 0.002560 loss= 8.824 step 9 lr= 0.005120 loss= 8.739 step 10 lr= 0.010240 loss= 8.570 step 11 lr= 0.020480 loss= 8.242 step 12 lr= 0.040960 loss= 7.619 step 13 lr= 0.081920 loss= 6.493 step 14 lr= 0.163840 loss= 4.668 step 15 lr= 0.327680 loss= 2.305 step 16 lr= 0.655360 loss= 0.453 step 17 lr= 1.310720 loss= 0.027 step 18 lr= 2.621440 loss= 0.059 step 19 lr= 5.242880 loss= 1.128 step 20 lr= 10.485760 loss= 131.190 step 21 lr= 20.971520 loss= 69417.688
A note on those last rows: once training crosses its breaking point, tiny floating-point differences amplify, so the blow-up digits can differ across torch versions and machines. The shape — steady, falling, rising, exploding — is what reproduces reliably.
Reading the curve
- Steps 0–8: loss barely moves. These rates are affordable but wasteful.
- Steps 13–17: loss falls hard. This is the productive zone.
- Steps 19–21: loss rises, then explodes. Past the cliff.
The standard advice: pick a rate in the steep part, roughly 3–10 times below the cliff. Here the cliff sits near 5, so something around 0.5–1 is aggressive and 0.1 is comfortable. If you plan to use one-cycle scheduling, the value you found becomes max_lr.
Walkthrough
Writing group["lr"] directly — optimisers read the learning rate fresh from their param groups at every step(), so a plain assignment is the whole mechanism. Schedulers do the same thing with more ceremony.
Doubling, not adding — learning rates live on a multiplicative scale. The interesting difference is between 0.001 and 0.01, not between 0.010 and 0.011. Real range tests grow the rate exponentially, as this one does.
One pass, growing within it — the sweep deliberately reuses one model. Each step is both a training step and a probe. That makes it cheap, and slightly dirty — see the mistakes below.
Common mistakes
Picking the rate at the lowest loss. The minimum of the curve sits at the edge of the cliff, and a schedule that lingers there diverges mid-training. Choose from the steep-descent region instead, below the minimum.
Reusing the swept model. By the end of the test, the model has been hit with enormous updates and its weights are damaged. Throw it away and retrain from scratch — the test's output is a number, not a model.
Trusting one seed. Runs this short are noisy. Two or three sweeps with different seeds cost little and stop you from tuning to a fluke. Seeding is covered in seeds and reproducibility.
Sweeping with the wrong batch size or optimiser. The good zone shifts with both. Rerun the test after switching from SGD to AdamW, or after changing batch size meaningfully — the old answer no longer applies.
Try it yourself
Replace SGD with torch.optim.AdamW (start lr=1e-6) and rerun. Find where the cliff moved, and compare the usable zone against SGD's. You will see with your own eyes why Adam defaults are 100x smaller.
What to learn next
- Learning rate schedulers — turning your found value into a full schedule.
- When the loss will not go down — diagnosing the case where no learning rate seems to help.
- Tracking down a NaN loss — what the far side of the cliff looks like in a real run.
Researcher — Mathematics and papers.
Origin
The LR range test is due to Smith (2017), Cyclical Learning Rates for Training Neural Networks — introduced as a by-product of cyclical schedules: sweep $\eta$ linearly or exponentially over a few epochs, plot loss (or accuracy) against $\eta$, and bound the cyclical schedule by the productive region. Smith and Topin (2019) reuse the test to select $\eta_{\max}$ for super-convergence.
What the cliff is
For a quadratic loss $\mathcal{L}(\theta) = \frac{1}{2}\theta^\top H \theta$, gradient descent with step $\eta$ has iteration matrix $(I - \eta H)$; convergence requires
$$\eta < \frac{2}{\lambda_{\max}(H)}$$
- $H$ — the Hessian, the matrix of second derivatives of the loss.
- $\lambda_{\max}$ — its largest eigenvalue, the curvature of the sharpest direction.
The explosion in the demo is this bound being crossed: for least squares, $H = \frac{2}{n}X^\top X$, fixed throughout training. For deep networks $H$ changes as weights move, so the usable rate drifts — one reason the test's answer is a starting zone, not a contract. Cohen et al. (2021), Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability, show that full-batch training tends to push $\lambda_{\max}$ up until it hovers at $2/\eta$, complicating the classical picture.
Batch size coupling
- Linear scaling rule (Goyal et al., 2017): multiply $\eta$ by $k$ when multiplying batch size by $k$, valid up to a problem-dependent limit; derived from treating $k$ small steps as one large one.
- For adaptive methods, square-root scaling is often argued instead (Krizhevsky, 2014, notes both; Malladi et al., 2022 give an SDE-based account).
- McCandlish et al. (2018), An Empirical Model of Large-Batch Training, define the gradient noise scale, predicting the batch size beyond which larger batches stop buying speed — the same regime where lr scaling breaks down.
Transfer across model size
Yang et al. (2021), Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer (muP) — under maximal-update parameterisation, the optimal learning rate becomes approximately width-invariant, so it can be tuned on a small proxy model and transferred. This, not per-run range tests, is how frontier-scale learning rates are now typically chosen.
Tooling
The torch-lr-finder package and Lightning's Tuner.lr_find implement the exponential sweep with smoothing and automatic suggestion (steepest negative gradient of the smoothed curve). Their suggestion heuristics differ; inspect the curve rather than accepting the marker blindly.
What to learn next
- Learning rate schedulers — turning your found value into a full schedule.
- When the loss will not go down — diagnosing the case where no learning rate seems to help.
- Tracking down a NaN loss — what the far side of the cliff looks like in a real run.