Optimisers, Schedulers and the Training Loop
Choosing between SGD, Adam and AdamW
SGD, Adam and AdamW are three strategies for turning gradients into weight updates, and each one changes how fast and how safely your model learns.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
An optimiser is the part of training that decides how to change each weight after the gradients are computed.
Picture three people walking down the same foggy hill. The first takes a careful step directly downhill, stops, and looks again. The second rides a bicycle, so speed built up on earlier slopes carries into the next stretch. The third wears smart boots that take long strides on flat, safe ground and short strides on rocky ground.
All three reach the bottom. They differ in how fast they get there, and how badly one wrong step hurts.
That is the difference between SGD (the careful stepper), SGD with momentum (the bicycle), and Adam (the smart boots). AdamW is Adam with one repair, which we will get to.
Why more than one optimiser exists
Gradients tell you which direction lowers the loss. They do not tell you how far to move.
Plain steps work, but they zigzag through narrow valleys and crawl across flat ground. Momentum smooths the zigzag by remembering the recent direction of travel. Adam goes further: it gives every single weight its own step size, tuned from that weight's recent history.
None of them wins everywhere. This is one of those honest cases where the field says "it depends" and means it.
How it works
gradients → [ optimiser + its memory of past steps ] → new weightsEvery optimiser sees the same gradients. The memory it keeps — none, a speed, or a per-weight history — is what makes each one behave differently.
Where you have seen this
Nearly every large language model you have used — the GPT family, Llama, Gemma — was trained with AdamW. Many famous image models were trained with SGD plus momentum. Both camps ship real products, which tells you both choices work.
Remember this
- The optimiser turns gradients into weight updates; each kind keeps different memory.
- Momentum smooths the path. Adam gives each weight its own step size.
- AdamW is today's default for transformers and fine-tuning.
What to learn next
- Parameter groups and per-layer learning rates — one optimiser, different rules for different layers.
- Learning rate schedulers — moving the learning rate while training runs.
- Gradient descent — the idea all three optimisers are built on.
Developer — Code and libraries.
Setup
pip install torchOutputs below were captured with torch 2.5.1 on CPU. Your last digits may differ slightly across versions and hardware.
The same race, four runners
Every run starts from identical weights, so the only difference is the optimiser.
import torch
import torch.nn as nn
torch.manual_seed(0)
X = torch.randn(256, 10)
y = (2 * X[:, 0] - X[:, 3] + 0.5 * X[:, 7] > 0).long() # a learnable rule with 3 useful features
def make_opt(name, params):
if name == "SGD": return torch.optim.SGD(params, lr=0.1)
if name == "SGD+momentum": return torch.optim.SGD(params, lr=0.1, momentum=0.9)
if name == "Adam": return torch.optim.Adam(params, lr=1e-3)
return torch.optim.AdamW(params, lr=1e-3, weight_decay=0.01)
loss_fn = nn.CrossEntropyLoss()
for name in ["SGD", "SGD+momentum", "Adam", "AdamW"]:
torch.manual_seed(1) # identical starting weights: a fair race
model = nn.Sequential(nn.Linear(10, 32), nn.ReLU(), nn.Linear(32, 2))
opt = make_opt(name, model.parameters())
for step in range(200):
opt.zero_grad()
loss = loss_fn(model(X), y)
loss.backward()
opt.step()
print(f"{name:12s} loss after 200 steps: {loss.item():.4f}")SGD loss after 200 steps: 0.0834 SGD+momentum loss after 200 steps: 0.0049 Adam loss after 200 steps: 0.2039 AdamW loss after 200 steps: 0.2042
Read that output carefully, because it teaches something most tutorials hide. Adam lost this race. Not because Adam is bad — because each optimiser ran at its own customary default learning rate. SGD ran at 0.1 and Adam at 0.001, the polite default from its paper.
The honest lesson: you cannot compare optimisers without tuning the learning rate for each one. An optimiser choice is a package deal with its learning rate.
Walkthrough
opt.zero_grad() — gradients accumulate by default in PyTorch. Skipping this line adds every previous step's gradients into the current one. See gradient accumulation for when that is done on purpose.
momentum=0.9 — keeps a running "velocity" per weight. Each update is mostly the previous update, plus a nudge from the new gradient. Loss went from 0.0834 to 0.0049 on the strength of this one argument.
Adam — keeps two running averages per weight: the recent gradient direction, and the recent gradient size. Weights with noisy, large gradients get careful small steps. Weights with steady, small gradients get bolder ones.
AdamW, weight_decay=0.01 — weight decay shrinks every weight slightly at each step, which discourages the model from leaning too hard on any single weight. Plain Adam mixes this shrinking into its per-weight machinery, which quietly weakens it. AdamW applies the shrink separately — the "W" stands for decoupled weight decay. When you want weight decay, prefer AdamW.
Which one should you pick?
| Situation | Reasonable default |
|---|---|
| Fine-tuning a transformer, or most modern work | AdamW, lr=1e-3 down to 1e-5, weight_decay=0.01 |
| Training a CNN from scratch, with time to tune | SGD, lr=0.1, momentum=0.9, plus a scheduler |
| No idea yet, want progress today | AdamW with lr=1e-3, then tune |
Adam-family optimisers forgive a poorly chosen learning rate more than SGD does. That forgiveness is why they dominate as defaults.
Common mistakes
Carrying a learning rate between optimisers. lr=0.1 is a normal SGD value and a catastrophic Adam value. Moving to Adam usually means dropping the learning rate by a factor of about 100.
Using weight_decay with plain Adam. It runs, but the decay gets scaled down for exactly the weights that update most. Switch to AdamW — the fix is one letter.
Recreating the optimiser every epoch. Building a fresh Adam inside the epoch loop wipes its running averages, so it never leaves its warm-up behaviour. Create it once, before the loop.
Forgetting momentum on SGD. torch.optim.SGD defaults to momentum=0. Nearly every published "SGD" result means SGD with momentum 0.9.
Try it yourself
Give Adam a fair fight: rerun the race with Adam(params, lr=1e-2). Then run SGD at lr=1e-3 and watch it crawl. Write down what this says about comparing optimisers at their defaults.
What to learn next
- Parameter groups and per-layer learning rates — one optimiser, different rules for different layers.
- Learning rate schedulers — moving the learning rate while training runs.
- Gradient descent — the idea all three optimisers are built on.
Researcher — Mathematics and papers.
Update rules
Let $g_t = \nabla_\theta \mathcal{L}(\theta_t)$ be the gradient at step $t$, and $\eta$ the learning rate.
SGD:
$$\theta_{t+1} = \theta_t - \eta\, g_t$$
SGD with momentum (Polyak, 1964):
$$v_{t+1} = \mu v_t + g_t, \qquad \theta_{t+1} = \theta_t - \eta\, v_{t+1}$$
- $v_t$ — the velocity, an exponentially-weighted sum of past gradients.
- $\mu$ — the momentum coefficient, typically $0.9$.
Adam (Kingma and Ba, 2015):
$$m_t = \beta_1 m_{t-1} + (1-\beta_1)\, g_t, \qquad v_t = \beta_2 v_{t-1} + (1-\beta_2)\, g_t^2$$
$$\hat{m}_t = \frac{m_t}{1-\beta_1^t}, \qquad \hat{v}_t = \frac{v_t}{1-\beta_2^t}, \qquad \theta_{t+1} = \theta_t - \eta\, \frac{\hat{m}_t}{\sqrt{\hat{v}_t} + \epsilon}$$
- $m_t$ — first-moment estimate (a running mean of gradients).
- $v_t$ — second-moment estimate (a running mean of squared gradients).
- $\beta_1, \beta_2$ — decay rates, defaults $0.9$ and $0.999$.
- $\hat{m}_t, \hat{v}_t$ — bias-corrected estimates; the correction matters early, when both averages start from zero.
- $\epsilon$ — numerical floor, default $10^{-8}$; some architectures are surprisingly sensitive to it.
AdamW (Loshchilov and Hutter, 2019) decouples weight decay from the adaptive machinery:
$$\theta_{t+1} = \theta_t - \eta \left( \frac{\hat{m}_t}{\sqrt{\hat{v}_t} + \epsilon} + \lambda\, \theta_t \right)$$
- $\lambda$ — the decay coefficient, applied directly to the weights. Folding L2 into the Adam gradient instead means parameters with large $\hat{v}_t$ receive almost no regularisation; decoupling restores uniform decay.
Cost
Per parameter: SGD stores no extra state, momentum stores one value, Adam and AdamW store two. For a 7B-parameter model in float32, Adam state adds roughly 56 GB on top of roughly 28 GB of weights — the motivation for 8-bit optimisers (Dettmers et al., 2022, 8-bit Optimizers via Block-wise Quantization) and paged or fused variants.
Convergence caveats and current practice
- Reddi et al. (2018), On the Convergence of Adam and Beyond — constructs problems where Adam provably fails to converge, and proposes AMSGrad. In practice, Adam with warm-up dominates anyway.
- Wilson et al. (2017), The Marginal Value of Adaptive Gradient Methods in Machine Learning — tuned SGD can out-generalise adaptive methods on vision tasks; the gap narrows with modern regularisation.
- Recent alternatives with real adoption: Lion (Chen et al., 2023), Sophia (Liu et al., 2023), and Muon (Jordan et al., 2024), which by 2025 shows strong results on transformer pre-training speedruns. AdamW remains the production default.
Papers
- Robbins and Monro (1951), A Stochastic Approximation Method — the ancestor of SGD.
- Kingma and Ba (2015), Adam: A Method for Stochastic Optimization.
- Loshchilov and Hutter (2019), Decoupled Weight Decay Regularization.
What to learn next
- Parameter groups and per-layer learning rates — one optimiser, different rules for different layers.
- Learning rate schedulers — moving the learning rate while training runs.
- Gradient descent — the idea all three optimisers are built on.