Weight decay
In one sentence Weight decay shrinks every weight slightly at each training step, so no weight grows large without earning it.
Updated
Weight decay shrinks every weight a little on every update, so weights only stay large if the data keeps proving they should.
It works like a small monthly fee on every subscription you hold. Services you truly use, you keep paying for. The ones you forgot about quietly drain away to nothing. Weight decay charges every parameter that fee: each training step multiplies weights by a number slightly below one, and only weights the gradients keep re-justifying stay big.
Why bother? Large weights let a network draw wild, sharp decision boundaries — the geometry of overfitting. Keeping weights small keeps the learned function smoother, which usually generalises better.
In practice you meet it as one argument in the optimizer:
torch.optim.AdamW(model.parameters(), lr=3e-4, weight_decay=0.01)One detail worth knowing: with plain SGD, weight decay and L2 regularization are equivalent. With Adam they are not, which is why AdamW exists — it applies the decay directly to the weights instead of mixing it into the gradient. AdamW is the default choice for training transformers today.
Where to go next
- Full lesson: Gradient descent
- Related terms: regularization, parameter, gradient-descent, overfitting