Optimisers, Schedulers and the Training Loop
Parameter groups and per-layer learning rates
One optimiser can treat different parts of a model differently — gentle updates for pretrained layers, bold updates for fresh ones — using parameter groups.
- 7 min read
- 3 reading levels
- Published
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Parameter groups let one optimiser apply different settings — a different learning rate, most often — to different parts of the same model.
Imagine restoring an old family portrait that has one torn corner. On the original painting, you work with feather-light touches, because it already looks right. On the fresh patch covering the tear, you paint boldly, because it starts blank.
One restorer, one brush, two very different pressures. That is a parameter group.
Why it exists
The usual reason is transfer learning — starting from a model someone already trained, and adapting it to your task. The borrowed layers already contain hard-won knowledge. The new final layer you bolted on contains random numbers.
Update both at the same speed and you face a bad trade. A speed suited to the new layer wrecks the borrowed knowledge. A speed suited to the borrowed layers leaves the new layer learning far too slowly.
Parameter groups end the trade: each part gets its own speed.
How it works
┌── group A: borrowed layers → tiny steps
model weights ──────┤
└── group B: new final layer → big steps
one optimiser, one step() call — two rulesYou hand the optimiser labelled bundles of weights instead of one big bag. Everything else about training stays the same.
Where you have seen this
Every "fine-tune this model on your own data" product works this way. Custom photo filters, voice models trained on your recordings, chatbots adapted to a company's documents — all of them. Under the hood, borrowed layers get gentle updates and new parts get bold ones.
Remember this
- A parameter group is a bundle of weights with its own optimiser settings.
- Standard use: pretrained parts move gently, new parts move fast.
- One optimiser, one
step()— the grouping does the rest.
What to learn next
- Learning rate schedulers — schedulers rewrite the group learning rates you set here.
- Transfer learning in PyTorch — the workflow that makes two-speed training necessary.
- requires_grad and leaf tensors — full freezing, when a slow group is not slow enough.
Developer — Code and libraries.
Setup
pip install torchOutputs captured with torch 2.5.1 on CPU. Exact decimals may vary across versions.
Two speeds in one optimiser
import torch
import torch.nn as nn
torch.manual_seed(0)
backbone = nn.Sequential(nn.Linear(8, 16), nn.ReLU()) # pretend this part is pretrained
head = nn.Linear(16, 2) # and this part is freshly added
model = nn.Sequential(backbone, head)
opt = torch.optim.AdamW([
{"params": backbone.parameters(), "lr": 1e-5}, # pretrained: nudge gently
{"params": head.parameters(), "lr": 1e-3}, # fresh: move fast
], weight_decay=0.01)
for i, group in enumerate(opt.param_groups):
n = sum(p.numel() for p in group["params"])
print(f"group {i}: lr={group['lr']:.0e} weight_decay={group['weight_decay']} params={n}")
before = [p.detach().clone() for p in model.parameters()]
loss = model(torch.randn(4, 8)).sum()
loss.backward()
opt.step()
moved = [(p - b).abs().max().item() for p, b in zip(model.parameters(), before)]
print(f"biggest backbone weight change: {max(moved[:2]):.6f}")
print(f"biggest head weight change: {max(moved[2:]):.6f}")group 0: lr=1e-05 weight_decay=0.01 params=144 group 1: lr=1e-03 weight_decay=0.01 params=34 biggest backbone weight change: 0.000010 biggest head weight change: 0.001002
One step() call, and the head moved about 100 times further than the backbone. The ratio of the changes matches the ratio of the learning rates.
Walkthrough
The list of dicts — instead of AdamW(model.parameters()), you pass a list. Each dict needs a "params" entry; any other key ("lr", "weight_decay", "betas") overrides the optimiser-wide default for that group only.
weight_decay=0.01 outside the list — settings passed after the list become the default for every group that does not override them. Both groups above inherited it, as the printout shows.
opt.param_groups — a live list you can read and write. Changing group["lr"] mid-training takes effect on the next step(). This is exactly how learning rate schedulers work under the hood.
The before/after measurement — comparing cloned weights around one step is a two-line experiment that proves the grouping worked. The same trick catches optimisers that update nothing, which is a real bug covered in when the loss will not go down.
A common real pattern: no decay for norms and biases
Transformer fine-tuning recipes usually exempt biases and LayerNorm weights from weight decay:
decay, no_decay = [], []
for name, p in model.named_parameters():
if p.ndim == 1: # biases and norm weights are 1-D
no_decay.append(p)
else:
decay.append(p)
opt = torch.optim.AdamW([
{"params": decay, "weight_decay": 0.01},
{"params": no_decay, "weight_decay": 0.0},
], lr=2e-5)Decaying a LayerNorm scale drags it toward zero, which fights what the layer is for. The 1-D test is the standard shortcut for finding these parameters.
Common mistakes
The same parameter in two groups. PyTorch raises ValueError: some parameters appear in more than one parameter group. It happens when you build groups from overlapping generators, such as model.parameters() in one group and head.parameters() in another.
Freezing when you meant "slow". lr=0 in a group still spends memory and time computing those gradients. To truly freeze layers, set requires_grad=False on them instead — see requires_grad and leaf tensors.
Exhausted generators. backbone.parameters() is a generator. Store list(backbone.parameters()) if you need it twice; a second pass over a used generator yields nothing, and the group ends up empty without an error.
Adding new layers after building the optimiser. Parameters created later are unknown to the optimiser and never update. Either rebuild the optimiser or call opt.add_param_group({"params": new_layer.parameters()}).
Try it yourself
Add a third group: split the backbone's two layers so the first gets lr=1e-6 and the second 1e-5. Confirm from the before/after measurement that the three parts move at three different scales.
What to learn next
- Learning rate schedulers — schedulers rewrite the group learning rates you set here.
- Transfer learning in PyTorch — the workflow that makes two-speed training necessary.
- requires_grad and leaf tensors — full freezing, when a slow group is not slow enough.
Researcher — Mathematics and papers.
Discriminative fine-tuning
The layerwise-rate idea was formalised for transfer learning by Howard and Ruder (2018), Universal Language Model Fine-tuning for Text Classification (ULMFiT), as discriminative fine-tuning:
$$\eta^{(l-1)} = \frac{\eta^{(l)}}{\xi}$$
- $\eta^{(l)}$ — the learning rate of layer $l$, with layers numbered from input to output.
- $\xi$ — the decay factor between adjacent layers; ULMFiT used $\xi = 2.6$.
The rationale: lower layers encode general features that transfer broadly and should move least, while upper layers are task-specific. For BERT-era encoders this reappears as layerwise learning-rate decay (LLRD) with $\xi \approx 1.05$–$1.3$, e.g. in Clark et al. (2020), ELECTRA, appendix hyperparameters.
Why norms and biases skip decay
Weight decay implements a zero-mean Gaussian prior over parameters. For LayerNorm gains (initialised at 1) and biases, that prior is centred in the wrong place, and empirically decaying them costs accuracy. The exemption appears in the original BERT fine-tuning code (Devlin et al., 2019) and has been default practice since.
Layer-adaptive methods
Per-layer scaling can be automated rather than hand-set:
- LARS (You et al., 2017, Large Batch Training of Convolutional Networks) — scales each layer's update by $|\theta^{(l)}| / |g^{(l)}|$, the trust ratio, enabling batch sizes in the tens of thousands.
- LAMB (You et al., 2020, Large Batch Optimization for Deep Learning) — the Adam-based counterpart, used to train BERT in 76 minutes.
Both are, mechanically, parameter groups whose learning rates are computed from layer statistics at every step.
Optimiser state is per-parameter, groups are per-setting
Adam's $m_t, v_t$ buffers attach to individual parameters, not to groups. Moving a parameter between groups (or rebuilding groups on checkpoint resume in a different order) silently reassociates hyperparameters while keeping state — a classic source of irreproducible resumes. optimizer.state_dict() serialises groups by index, so group construction order must be stable across save and load.
What to learn next
- Learning rate schedulers — schedulers rewrite the group learning rates you set here.
- Transfer learning in PyTorch — the workflow that makes two-speed training necessary.
- requires_grad and leaf tensors — full freezing, when a slow group is not slow enough.