Optimisers, Schedulers and the Training Loop

Learning rate schedulers

A scheduler changes the learning rate while training runs — fast at the start, gentle at the end — and PyTorch ships the common shapes ready-made.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have seen this
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A learning rate scheduler automatically changes the size of the training steps as training goes on.

Think about parking a scooter. Far from the wall you roll in quickly, because precision does not matter yet. Close to the wall you slow to a crawl, because now a big move means a dent.

Training works the same way. Early on, the model is far from good settings, so big steps save time. Near the end, big steps overshoot the good settings and undo progress.

Why it exists

One fixed step size is a bad deal both ways. A size that is safe for the end wastes hours at the start. A size that is fast for the start ruins the finish.

People used to lower the step size by hand when training looked stuck. A scheduler is that habit, automated: a rule that moves the learning rate along a planned curve.

How it works

learning
rate      ████
          █   ███
          █      ███
          █         ████
          █             █████
          └────────────────────────  training time

fast and rough at first  →  slow and careful at the end

Different schedules draw different curves — stair steps, a smooth wave, or a quick ramp up followed by a long glide down. The shape varies; the idea is the same.

Where you have seen this

Every headline model trains with a scheduled learning rate: GPT-style language models, image generators, translation systems. The usual shape is a short warm-up followed by a long smooth decay. A fixed rate from start to finish is now rare outside quick experiments.

Remember this

  • The learning rate should usually shrink as training progresses.
  • A scheduler automates that shrinking along a planned curve.
  • Warm-up, stair steps and smooth cosine glides are the common shapes.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Outputs captured with torch 2.5.1 on CPU; values are deterministic.

Three schedules, printed honestly

schedulers.py
import torch
import torch.nn as nn

def lr_history(make_sched, epochs=10):
    model = nn.Linear(4, 2)
    opt = torch.optim.SGD(model.parameters(), lr=0.1)
    sched = make_sched(opt)
    lrs = []
    for _ in range(epochs):
        opt.step()          # training for one epoch would happen here
        sched.step()        # then the scheduler moves the lr
        lrs.append(sched.get_last_lr()[0])
    return lrs

S = torch.optim.lr_scheduler
runs = {
    "StepLR(step_size=3, gamma=0.1)": lambda o: S.StepLR(o, step_size=3, gamma=0.1),
    "CosineAnnealingLR(T_max=10)":    lambda o: S.CosineAnnealingLR(o, T_max=10),
    "LinearLR(0.1 -> 1.0, 5 epochs)": lambda o: S.LinearLR(o, start_factor=0.1, total_iters=5),
}
for name, make in runs.items():
    print(name)
    print("  " + "  ".join(f"{lr:.4f}" for lr in lr_history(make)))
Output
StepLR(step_size=3, gamma=0.1)
  0.1000  0.1000  0.0100  0.0100  0.0100  0.0010  0.0010  0.0010  0.0001  0.0001
CosineAnnealingLR(T_max=10)
  0.0976  0.0905  0.0794  0.0655  0.0500  0.0345  0.0206  0.0095  0.0024  0.0000
LinearLR(0.1 -> 1.0, 5 epochs)
  0.0280  0.0460  0.0640  0.0820  0.1000  0.1000  0.1000  0.1000  0.1000  0.1000

Three personalities. StepLR drops in stairs — ten times smaller every three epochs. CosineAnnealingLR glides down a smooth wave to zero across T_max epochs. LinearLR is the odd one: it goes up, ramping from 10% of the base rate to full strength — that ramp is warm-up, a gentle start that protects fragile early training.

Walkthrough

The scheduler wraps the optimiser, not the model. It works by rewriting lr inside opt.param_groups — the same mechanism you used by hand in parameter groups. With several groups, each is scaled from its own base rate.

opt.step() before sched.step(). That order is required. Reverse it and recent PyTorch versions warn that the first scheduler step ran before any optimiser step, skewing the schedule by one.

get_last_lr() — returns a list, one value per parameter group, hence the [0].

How often to call step() depends on the scheduler. StepLR and CosineAnnealingLR are designed for once per epoch. OneCycleLR — a popular ramp-up-then-glide schedule built for a fixed number of steps — expects a call after every batch and needs total_steps (or epochs and steps_per_epoch) at construction. Warm-up ramps like LinearLR are also usually stepped per batch.

Combining warm-up with a decay — chain them:

python
sched = S.SequentialLR(opt,
    schedulers=[S.LinearLR(opt, start_factor=0.1, total_iters=5),
                S.CosineAnnealingLR(opt, T_max=95)],
    milestones=[5])

Five epochs of ramp, then a 95-epoch cosine glide — the standard modern recipe.

Common mistakes

Stepping a per-epoch scheduler every batch. CosineAnnealingLR(T_max=10) stepped per batch finishes its entire glide during epoch one, and the model spends the rest of training at a learning rate near zero — training that mysteriously flatlines early. Match the stepping frequency to the scheduler's design.

Forgetting the scheduler in checkpoints. Schedulers have state_dict() and load_state_dict() like models. Resuming a run without restoring the scheduler restarts the curve from the top, hitting your half-trained model with a full-strength learning rate.

Assuming ReduceLROnPlateau behaves like the others. It reacts to a metric rather than a clock, so its call signature is sched.step(val_loss). Calling it with no argument raises a TypeError.

Tuning the base rate while ignoring the floor. Cosine schedules end at eta_min, default 0. If training a long tail matters, set eta_min=1e-5 or similar; a literal zero learning rate makes the final epochs decorative.

Try it yourself

Add S.OneCycleLR(o, max_lr=0.1, total_steps=10) to the runs dict and print its ten values. Identify where the peak lands and what the final rate is, then compare the shape to the cosine run.

What to learn next

Researcher — Mathematics and papers.

Cosine annealing

Loshchilov and Hutter (2017), SGDR: Stochastic Gradient Descent with Warm Restarts, set the schedule

$$\eta_t = \eta_{\min} + \frac{1}{2}\left(\eta_{\max} - \eta_{\min}\right)\left(1 + \cos!\left(\frac{T_{cur}}{T_{max}}\,\pi\right)\right)$$

  • $\eta_{\max}$ — the base learning rate; $\eta_{\min}$ — the floor (PyTorch default 0).
  • $T_{cur}$ — epochs since the last restart; $T_{max}$ — the period length.

The restart half of SGDR (periodically resetting $T_{cur}$) is available as CosineAnnealingWarmRestarts but is used far less than the single-cycle version.

Warm-up: why the start is fragile

Two established motivations:

  • Large-batch stability. Goyal et al. (2017), Accurate, Large Minibatch SGD, pair the linear scaling rule with a linear warm-up over 5 epochs, preventing divergence at high effective rates.
  • Adam's early variance problem. Liu et al. (2020), On the Variance of the Adaptive Learning Rate and Beyond, show that Adam's $\hat{v}_t$ estimate has excessive variance in the first steps, making early updates erratically large; warm-up masks this, and their RAdam rectifies it analytically.

Transformers are notably warm-up-dependent: Vaswani et al. (2017) used the inverse-square-root schedule $\eta_t \propto \min(t^{-1/2},\, t \cdot T_w^{-3/2})$ with $T_w$ warm-up steps, and training BERT-scale models without warm-up frequently diverges (Xiong et al., 2020, On Layer Normalization in the Transformer Architecture, analyse why Post-LN makes this worse).

One-cycle and super-convergence

Smith and Topin (2019), Super-Convergence: Very Fast Training of Neural Networks Using Large Learning Rates, ramp up to a deliberately high $\eta_{\max}$ and back down within a single cycle, with momentum cycled inversely. The claimed regularisation effect of the high-rate phase lets some models train in far fewer epochs. OneCycleLR implements this including the momentum cycling.

Current LLM practice

Large-model pre-training now favours warmup-stable-decay (WSD) shapes: warm-up, a long constant plateau, then a sharp decay — attractive because the decay point can be chosen late (Hu et al., 2024, MiniCPM; Hägele et al., 2024 analyse it against cosine). Cosine-to-10%-of-peak remains the most common published recipe (e.g. Chinchilla, Hoffmann et al., 2022; Llama, Touvron et al., 2023).

What to learn next