Optimisers, Schedulers and the Training Loop

Finding a learning rate that works

The LR range test raises the learning rate step by step until training breaks, and the curve it draws tells you where a good value lives.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have seen this
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The fastest way to find a good learning rate is to raise it gradually until training breaks, then settle a comfortable distance below the breaking point.

It works like adjusting an unfamiliar shower tap. You turn the knob a little — nothing changes. A little more — pleasantly warm. Keep turning and suddenly it scalds.

Nobody finds the right setting by guessing one position and stepping in. You sweep the knob, feel where things go wrong, and settle safely below that point.

Why it exists

The learning rate is the single most influential setting in training. Too small and a model that could finish overnight takes weeks. Too large and the model never improves at all — it thrashes, or breaks completely.

The painful part: the good range differs between models, datasets and optimisers. There is no universal number. So instead of a number, you learn a procedure that finds the range in one cheap run.

How it works

loss
 │████                                  ██
 │    ████                            ██
 │        ██████                    ██
 │              ██████            ██
 │                    ████████████
 └──────────────────────────────────────  learning rate (growing)
   too small:      good zone:      too big:
   nothing moves   loss falls      loss explodes

Run a short training where the learning rate grows every step. Watch the loss. It stays flat, then falls, then explodes. The good values sit in the falling zone, safely left of the explosion.

Where you have seen this

The button-press version ships in real tools: the fastai library made "find my learning rate" a one-line command, and PyTorch Lightning includes an automatic tuner. Both run exactly this sweep behind the scenes.

Remember this

  • The learning rate is the most important knob in training.
  • Sweep it upward in one cheap run; the loss curve shows the good zone.
  • Pick a value in the steep-descent part, well below the explosion.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Outputs captured with torch 2.5.1 on CPU; this run is seeded and deterministic.

The range test, hand-rolled

Twenty-two steps, doubling the learning rate each time:

lr_range_test.py
import torch
import torch.nn as nn

torch.manual_seed(0)
X = torch.randn(512, 20)
true_w = torch.randn(20, 1)
y = X @ true_w + 0.1 * torch.randn(512, 1)

model = nn.Linear(20, 1)
opt = torch.optim.SGD(model.parameters(), lr=1e-5)
loss_fn = nn.MSELoss()

lr = 1e-5
for step in range(22):
    for group in opt.param_groups:
        group["lr"] = lr                     # raise the lr a little every step
    opt.zero_grad()
    loss = loss_fn(model(X), y)
    loss.backward()
    opt.step()
    print(f"step {step:2d}  lr={lr:10.6f}  loss={loss.item():14.3f}")
    lr *= 2.0                                # double it and see what breaks
Output
step  0  lr=  0.000010  loss=         8.909
step  1  lr=  0.000020  loss=         8.909
step  2  lr=  0.000040  loss=         8.908
step  3  lr=  0.000080  loss=         8.907
step  4  lr=  0.000160  loss=         8.904
step  5  lr=  0.000320  loss=         8.899
step  6  lr=  0.000640  loss=         8.888
step  7  lr=  0.001280  loss=         8.867
step  8  lr=  0.002560  loss=         8.824
step  9  lr=  0.005120  loss=         8.739
step 10  lr=  0.010240  loss=         8.570
step 11  lr=  0.020480  loss=         8.242
step 12  lr=  0.040960  loss=         7.619
step 13  lr=  0.081920  loss=         6.493
step 14  lr=  0.163840  loss=         4.668
step 15  lr=  0.327680  loss=         2.305
step 16  lr=  0.655360  loss=         0.453
step 17  lr=  1.310720  loss=         0.027
step 18  lr=  2.621440  loss=         0.059
step 19  lr=  5.242880  loss=         1.128
step 20  lr= 10.485760  loss=       131.190
step 21  lr= 20.971520  loss=     69417.688

A note on those last rows: once training crosses its breaking point, tiny floating-point differences amplify, so the blow-up digits can differ across torch versions and machines. The shape — steady, falling, rising, exploding — is what reproduces reliably.

Reading the curve

  • Steps 0–8: loss barely moves. These rates are affordable but wasteful.
  • Steps 13–17: loss falls hard. This is the productive zone.
  • Steps 19–21: loss rises, then explodes. Past the cliff.

The standard advice: pick a rate in the steep part, roughly 3–10 times below the cliff. Here the cliff sits near 5, so something around 0.5–1 is aggressive and 0.1 is comfortable. If you plan to use one-cycle scheduling, the value you found becomes max_lr.

Walkthrough

Writing group["lr"] directly — optimisers read the learning rate fresh from their param groups at every step(), so a plain assignment is the whole mechanism. Schedulers do the same thing with more ceremony.

Doubling, not adding — learning rates live on a multiplicative scale. The interesting difference is between 0.001 and 0.01, not between 0.010 and 0.011. Real range tests grow the rate exponentially, as this one does.

One pass, growing within it — the sweep deliberately reuses one model. Each step is both a training step and a probe. That makes it cheap, and slightly dirty — see the mistakes below.

Common mistakes

Picking the rate at the lowest loss. The minimum of the curve sits at the edge of the cliff, and a schedule that lingers there diverges mid-training. Choose from the steep-descent region instead, below the minimum.

Reusing the swept model. By the end of the test, the model has been hit with enormous updates and its weights are damaged. Throw it away and retrain from scratch — the test's output is a number, not a model.

Trusting one seed. Runs this short are noisy. Two or three sweeps with different seeds cost little and stop you from tuning to a fluke. Seeding is covered in seeds and reproducibility.

Sweeping with the wrong batch size or optimiser. The good zone shifts with both. Rerun the test after switching from SGD to AdamW, or after changing batch size meaningfully — the old answer no longer applies.

Try it yourself

Replace SGD with torch.optim.AdamW (start lr=1e-6) and rerun. Find where the cliff moved, and compare the usable zone against SGD's. You will see with your own eyes why Adam defaults are 100x smaller.

What to learn next

Researcher — Mathematics and papers.

Origin

The LR range test is due to Smith (2017), Cyclical Learning Rates for Training Neural Networks — introduced as a by-product of cyclical schedules: sweep $\eta$ linearly or exponentially over a few epochs, plot loss (or accuracy) against $\eta$, and bound the cyclical schedule by the productive region. Smith and Topin (2019) reuse the test to select $\eta_{\max}$ for super-convergence.

What the cliff is

For a quadratic loss $\mathcal{L}(\theta) = \frac{1}{2}\theta^\top H \theta$, gradient descent with step $\eta$ has iteration matrix $(I - \eta H)$; convergence requires

$$\eta < \frac{2}{\lambda_{\max}(H)}$$

  • $H$ — the Hessian, the matrix of second derivatives of the loss.
  • $\lambda_{\max}$ — its largest eigenvalue, the curvature of the sharpest direction.

The explosion in the demo is this bound being crossed: for least squares, $H = \frac{2}{n}X^\top X$, fixed throughout training. For deep networks $H$ changes as weights move, so the usable rate drifts — one reason the test's answer is a starting zone, not a contract. Cohen et al. (2021), Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability, show that full-batch training tends to push $\lambda_{\max}$ up until it hovers at $2/\eta$, complicating the classical picture.

Batch size coupling

  • Linear scaling rule (Goyal et al., 2017): multiply $\eta$ by $k$ when multiplying batch size by $k$, valid up to a problem-dependent limit; derived from treating $k$ small steps as one large one.
  • For adaptive methods, square-root scaling is often argued instead (Krizhevsky, 2014, notes both; Malladi et al., 2022 give an SDE-based account).
  • McCandlish et al. (2018), An Empirical Model of Large-Batch Training, define the gradient noise scale, predicting the batch size beyond which larger batches stop buying speed — the same regime where lr scaling breaks down.

Transfer across model size

Yang et al. (2021), Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer (muP) — under maximal-update parameterisation, the optimal learning rate becomes approximately width-invariant, so it can be tuned on a small proxy model and transferred. This, not per-run range tests, is how frontier-scale learning rates are now typically chosen.

Tooling

The torch-lr-finder package and Lightning's Tuner.lr_find implement the exponential sweep with smoothing and automatic suggestion (steepest negative gradient of the smoothed curve). Their suggestion heuristics differ; inspect the curve rather than accepting the marker blindly.

What to learn next