Gradient clipping
Clipping caps how large a gradient is allowed to get before the update, turning the occasional explosive step into a survivable one — same direction, sane size.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Gradient clipping puts a ceiling on how big an update step can be, so one bad batch cannot fling the model into nonsense.
Think of a governor on a school-van engine — the device that stops it exceeding a set speed. The driver still steers wherever the road goes; the van still moves the direction it was asked. It only cannot go dangerously fast, no matter how hard the pedal is pressed. Clipping is that governor for learning.
Why it exists
Sometimes a batch produces an enormous gradient — a rare weird sample, an unstable stretch of training, a deep network multiplying blame layer after layer until it snowballs. This snowballing has a name: exploding gradients.
One giant step can undo a week of training in a single update — the loss suddenly reads NaN (not-a-number: the arithmetic has broken down) and the model is effectively dead. Clipping does not fix the cause of the giant gradient. It makes the giant survivable, which keeps training alive long enough to proceed.
How it works
The standard version measures the overall size of all gradients together, and shrinks them proportionally when the size exceeds a limit:
gradient size 3 , limit 10 -> untouched
gradient size 250 , limit 10 -> every gradient scaled down by 25x
direction kept, size now exactly 10Direction preserved, size capped. The model still learns from the bad batch — modestly instead of catastrophically.
A real example you have seen
Voice assistants and translation apps are built on the network families where explosions were once routine — the ones reading sequences word by word. Clipping was the workhorse fix that made training them dependable, and it remains standard practice in the recipes behind today's chat models.
Remember this
- Exploding gradients: blame snowballs into giant, training-destroying steps.
- Clipping caps the step's size while keeping its direction.
- It treats the symptom well — and that is often exactly what training needs.
What to learn next
- Writing a custom autograd Function — going below the built-in operations.
- Gradient accumulation for large batches — the technique clipping must be sequenced with.
- LSTM — the architecture born from the vanishing half of this story.
Developer — Code and libraries.
Setup
pip install torchOutputs verified with torch 2.5.1, CPU.
Watching a clip happen
import torch
w = torch.tensor([3., 4.], requires_grad=True)
loss = (w * torch.tensor([30., 40.])).sum()
loss.backward()
print(w.grad, "norm =", w.grad.norm().item())
before = torch.nn.utils.clip_grad_norm_([w], max_norm=1.0)
print("returned:", before.item()) # the norm BEFORE clipping
print(w.grad, "norm =", w.grad.norm().item()) # same direction, length 1tensor([30., 40.]) norm = 50.0 returned: 50.0 tensor([0.6000, 0.8000]) norm = 0.9999999403953552
The walkthrough
The gradient [30, 40] has overall size 50 (the 3-4-5 triangle, scaled). With a limit of 1.0, every element is multiplied by 1/50: direction identical, size 1. That final 0.99999994 instead of 1.0 is ordinary float32 rounding — worth seeing once so it never worries you.
The return value is the norm before clipping. Log it every step. That single number is your explosion seismograph: healthy training shows it wandering within a band; trouble shows spikes — and the log tells you when they started, which batch, and whether your limit is doing anything at all.
Placement is rigid — after backward(), before step(), on the completed gradients:
loss.backward()
grad_norm = torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
optimizer.step()
optimizer.zero_grad()(Sketch — slots into any real loop.) With gradient accumulation, clip once per update on the finished pile, never per micro-batch.
The blunt alternative is clipping each element separately to a range:
import torch
w = torch.tensor([3., 4.], requires_grad=True)
(w * torch.tensor([30., 40.])).sum().backward()
torch.nn.utils.clip_grad_value_([w], clip_value=5.0)
print(w.grad)tensor([5., 5.])
Note what happened: [30, 40] became [5, 5] — the direction changed (elements hit the cap unequally). Norm clipping is the default choice in modern recipes precisely because it preserves direction; value clipping survives in some reinforcement-learning code.
Choosing the limit
Common starting points: 1.0 (transformer training), 0.5–5.0 elsewhere. Better than folklore: run without clipping for a while, log gradient norms, and set the limit around the 90th–95th percentile of healthy values. The governor should catch runaways, not lean on every step — a limit that clips constantly is quietly shrinking your effective learning rate.
Common mistakes
Clipping before backward, or after step. Before backward there is nothing to clip; after step the damage is applied. The order in loop_pattern.py is the only correct one.
Using clipping as a NaN cure. If your loss is already NaN, the gradients are NaN, and scaling NaN gives NaN — clipping cannot resurrect it. Clipping prevents explosions from feeding forward; existing NaNs mean a different bug (learning rate, a log of zero, division by zero). Fix the cause, then clip to prevent recurrence.
Clipping parameters instead of gradients. clip_grad_norm_ touches .grad only. Clamping weights is a different (rarely appropriate) intervention. The function name says grad; trust it.
Setting the limit brutally low to "be safe". Norms permanently above the limit mean every update is rescaled — training slows, and your learning-rate tuning is fighting the governor. The logged pre-clip norm tells you when this is happening.
Forgetting the mixed-precision order. With GradScaler, gradients are scaled up; clip them un-scaled — scaler.unscale_(optimizer) first, then clip, then scaler.step. Clipping scaled gradients applies a meaningless threshold.
Try it yourself
Log clip_grad_norm_'s return value while training any small model for 200 steps, then plot the series. Find your 95th percentile and compare it with the limit you had guessed. Then multiply your learning rate by 10 and watch what the seismograph does — explosions on demand, safely.
What to learn next
- Writing a custom autograd Function — going below the built-in operations.
- Gradient accumulation for large batches — the technique clipping must be sequenced with.
- LSTM — the architecture born from the vanishing half of this story.
Researcher — Mathematics and papers.
Definitions
Global norm clipping over parameter set {θ_i} with threshold c:
g ← g · min(1, c / (‖g‖₂ + ε))
Where g is the concatenation of all parameter gradients, ‖g‖₂ its global L2 norm (computed across all parameters jointly — the default in clip_grad_norm_), and ε a small constant against division by zero. Direction is preserved exactly; magnitude is capped at c. Value clipping applies g_i ← max(−c, min(c, g_i)) per element and is not direction-preserving.
Cost: one pass to compute the norm, one to scale — O(P) in parameter count, negligible against the backward pass. clip_grad_norm_ also accepts norm_type (any p-norm, including inf) and foreach=True fused implementations in torch 2.x.
Why gradients explode
For a depth-T composition (an unrolled RNN over T steps, or a T-layer net), the Jacobian of early state with respect to late loss contains products Π_{t} J_t. If the spectral radii of the J_t sit above 1, the product grows exponentially in T; below 1, it vanishes. This analysis — and clipping as the practical response to the exploding half — is Pascanu, Mikolov and Bengio (2013), On the difficulty of training recurrent neural networks, ICML. The vanishing half required architectural fixes instead: gating in LSTMs, and residual connections in deep feedforward nets.
Theoretical footing beyond folklore: Zhang et al. (2020), Why Gradient Clipping Accelerates Training, ICLR, prove convergence guarantees for clipped GD under a relaxed smoothness condition (L₀, L₁-smoothness: local smoothness growing with gradient norm) that better matches measured loss landscapes of deep nets than the classical Lipschitz assumption — under it, clipped GD provably beats any fixed-step GD. This reframes clipping from a hack into the adaptive step-size rule matching the geometry.
Relatives and variants
- Adaptive clipping: thresholds tracked from gradient-norm history (e.g. percentile rules; AutoClip, Seetharaman et al., 2020). AGC in NFNets (Brock et al., 2021) clips per-unit by the ratio ‖g‖/‖w‖, removing the global scale sensitivity.
- Per-sample clipping in DP-SGD (Abadi et al., 2016, Deep Learning with Differential Privacy): each sample's gradient clipped individually before aggregation — there the clip bounds sensitivity for the privacy accountant, a different purpose sharing the formula.
- Normalised/sign methods: normalised SGD and signSGD are the c → 0 limit family — direction-only updates; clipping interpolates between raw and normalised gradients.
- Under distributed training, clip after gradient synchronisation (DDP hooks handle ordering), so all replicas clip the same global norm.
References
- Pascanu, Mikolov, Bengio (2013), arXiv:1211.5063 — the exploding/vanishing analysis and the norm-clipping prescription.
- Zhang, He, Sra, Jadbabaie (2020), arXiv:1905.11881 — convergence theory for clipping.
- Abadi et al. (2016), CCS — per-sample clipping for privacy.
- Brock et al. (2021), arXiv:2102.06171 — adaptive gradient clipping in NFNets.
What to learn next
- Writing a custom autograd Function — going below the built-in operations.
- Gradient accumulation for large batches — the technique clipping must be sequenced with.
- LSTM — the architecture born from the vanishing half of this story.