Autograd in Depth

Gradient clipping

Clipping caps how large a gradient is allowed to get before the update, turning the occasional explosive step into a survivable one — same direction, sane size.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Gradient clipping puts a ceiling on how big an update step can be, so one bad batch cannot fling the model into nonsense.

Think of a governor on a school-van engine — the device that stops it exceeding a set speed. The driver still steers wherever the road goes; the van still moves the direction it was asked. It only cannot go dangerously fast, no matter how hard the pedal is pressed. Clipping is that governor for learning.

Why it exists

Sometimes a batch produces an enormous gradient — a rare weird sample, an unstable stretch of training, a deep network multiplying blame layer after layer until it snowballs. This snowballing has a name: exploding gradients.

One giant step can undo a week of training in a single update — the loss suddenly reads NaN (not-a-number: the arithmetic has broken down) and the model is effectively dead. Clipping does not fix the cause of the giant gradient. It makes the giant survivable, which keeps training alive long enough to proceed.

How it works

The standard version measures the overall size of all gradients together, and shrinks them proportionally when the size exceeds a limit:

gradient size 3   , limit 10   ->  untouched
gradient size 250 , limit 10   ->  every gradient scaled down by 25x
                                   direction kept, size now exactly 10

Direction preserved, size capped. The model still learns from the bad batch — modestly instead of catastrophically.

A real example you have seen

Voice assistants and translation apps are built on the network families where explosions were once routine — the ones reading sequences word by word. Clipping was the workhorse fix that made training them dependable, and it remains standard practice in the recipes behind today's chat models.

Remember this

  • Exploding gradients: blame snowballs into giant, training-destroying steps.
  • Clipping caps the step's size while keeping its direction.
  • It treats the symptom well — and that is often exactly what training needs.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Outputs verified with torch 2.5.1, CPU.

Watching a clip happen

clip.py
import torch

w = torch.tensor([3., 4.], requires_grad=True)
loss = (w * torch.tensor([30., 40.])).sum()
loss.backward()
print(w.grad, "norm =", w.grad.norm().item())

before = torch.nn.utils.clip_grad_norm_([w], max_norm=1.0)
print("returned:", before.item())               # the norm BEFORE clipping
print(w.grad, "norm =", w.grad.norm().item())   # same direction, length 1
Output
tensor([30., 40.]) norm = 50.0
returned: 50.0
tensor([0.6000, 0.8000]) norm = 0.9999999403953552

The walkthrough

The gradient [30, 40] has overall size 50 (the 3-4-5 triangle, scaled). With a limit of 1.0, every element is multiplied by 1/50: direction identical, size 1. That final 0.99999994 instead of 1.0 is ordinary float32 rounding — worth seeing once so it never worries you.

The return value is the norm before clipping. Log it every step. That single number is your explosion seismograph: healthy training shows it wandering within a band; trouble shows spikes — and the log tells you when they started, which batch, and whether your limit is doing anything at all.

Placement is rigid — after backward(), before step(), on the completed gradients:

loop_pattern.py
loss.backward()
grad_norm = torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
optimizer.step()
optimizer.zero_grad()

(Sketch — slots into any real loop.) With gradient accumulation, clip once per update on the finished pile, never per micro-batch.

The blunt alternative is clipping each element separately to a range:

python
import torch
w = torch.tensor([3., 4.], requires_grad=True)
(w * torch.tensor([30., 40.])).sum().backward()
torch.nn.utils.clip_grad_value_([w], clip_value=5.0)
print(w.grad)
Output
tensor([5., 5.])

Note what happened: [30, 40] became [5, 5] — the direction changed (elements hit the cap unequally). Norm clipping is the default choice in modern recipes precisely because it preserves direction; value clipping survives in some reinforcement-learning code.

Choosing the limit

Common starting points: 1.0 (transformer training), 0.5–5.0 elsewhere. Better than folklore: run without clipping for a while, log gradient norms, and set the limit around the 90th–95th percentile of healthy values. The governor should catch runaways, not lean on every step — a limit that clips constantly is quietly shrinking your effective learning rate.

Common mistakes

Clipping before backward, or after step. Before backward there is nothing to clip; after step the damage is applied. The order in loop_pattern.py is the only correct one.

Using clipping as a NaN cure. If your loss is already NaN, the gradients are NaN, and scaling NaN gives NaN — clipping cannot resurrect it. Clipping prevents explosions from feeding forward; existing NaNs mean a different bug (learning rate, a log of zero, division by zero). Fix the cause, then clip to prevent recurrence.

Clipping parameters instead of gradients. clip_grad_norm_ touches .grad only. Clamping weights is a different (rarely appropriate) intervention. The function name says grad; trust it.

Setting the limit brutally low to "be safe". Norms permanently above the limit mean every update is rescaled — training slows, and your learning-rate tuning is fighting the governor. The logged pre-clip norm tells you when this is happening.

Forgetting the mixed-precision order. With GradScaler, gradients are scaled up; clip them un-scaled — scaler.unscale_(optimizer) first, then clip, then scaler.step. Clipping scaled gradients applies a meaningless threshold.

Try it yourself

Log clip_grad_norm_'s return value while training any small model for 200 steps, then plot the series. Find your 95th percentile and compare it with the limit you had guessed. Then multiply your learning rate by 10 and watch what the seismograph does — explosions on demand, safely.

What to learn next

Researcher — Mathematics and papers.

Definitions

Global norm clipping over parameter set {θ_i} with threshold c:

g ← g · min(1, c / (‖g‖₂ + ε))

Where g is the concatenation of all parameter gradients, ‖g‖₂ its global L2 norm (computed across all parameters jointly — the default in clip_grad_norm_), and ε a small constant against division by zero. Direction is preserved exactly; magnitude is capped at c. Value clipping applies g_i ← max(−c, min(c, g_i)) per element and is not direction-preserving.

Cost: one pass to compute the norm, one to scale — O(P) in parameter count, negligible against the backward pass. clip_grad_norm_ also accepts norm_type (any p-norm, including inf) and foreach=True fused implementations in torch 2.x.

Why gradients explode

For a depth-T composition (an unrolled RNN over T steps, or a T-layer net), the Jacobian of early state with respect to late loss contains products Π_{t} J_t. If the spectral radii of the J_t sit above 1, the product grows exponentially in T; below 1, it vanishes. This analysis — and clipping as the practical response to the exploding half — is Pascanu, Mikolov and Bengio (2013), On the difficulty of training recurrent neural networks, ICML. The vanishing half required architectural fixes instead: gating in LSTMs, and residual connections in deep feedforward nets.

Theoretical footing beyond folklore: Zhang et al. (2020), Why Gradient Clipping Accelerates Training, ICLR, prove convergence guarantees for clipped GD under a relaxed smoothness condition (L₀, L₁-smoothness: local smoothness growing with gradient norm) that better matches measured loss landscapes of deep nets than the classical Lipschitz assumption — under it, clipped GD provably beats any fixed-step GD. This reframes clipping from a hack into the adaptive step-size rule matching the geometry.

Relatives and variants

  • Adaptive clipping: thresholds tracked from gradient-norm history (e.g. percentile rules; AutoClip, Seetharaman et al., 2020). AGC in NFNets (Brock et al., 2021) clips per-unit by the ratio ‖g‖/‖w‖, removing the global scale sensitivity.
  • Per-sample clipping in DP-SGD (Abadi et al., 2016, Deep Learning with Differential Privacy): each sample's gradient clipped individually before aggregation — there the clip bounds sensitivity for the privacy accountant, a different purpose sharing the formula.
  • Normalised/sign methods: normalised SGD and signSGD are the c → 0 limit family — direction-only updates; clipping interpolates between raw and normalised gradients.
  • Under distributed training, clip after gradient synchronisation (DDP hooks handle ordering), so all replicas clip the same global norm.

References

  • Pascanu, Mikolov, Bengio (2013), arXiv:1211.5063 — the exploding/vanishing analysis and the norm-clipping prescription.
  • Zhang, He, Sra, Jadbabaie (2020), arXiv:1905.11881 — convergence theory for clipping.
  • Abadi et al. (2016), CCS — per-sample clipping for privacy.
  • Brock et al. (2021), arXiv:2102.06171 — adaptive gradient clipping in NFNets.

What to learn next