Debugging PyTorch

Checking that gradients reach every layer

One loop over named_parameters after backward shows which layers received gradients, which got none, and which got signals too small to matter.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have seen this
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

After every backward pass, each layer should have received its share of the learning signal — and you can inspect, layer by layer, whether it actually did.

Picture a row of plants watered through one long pipe with an outlet at each plant. If a joint in the middle is blocked, every plant behind the blockage stays dry — while the plants before it flourish. The gardener's diagnostic is beautifully direct: walk the row and touch the soil at each plant.

In training, the water is the gradient — the correction signal flowing backward from the loss through every layer. A blocked pipe means some layers never receive corrections, never change, and never learn — while the rest of the model trains normally around them.

Why it exists

Partial blockages are among the sneakiest training bugs. The model does learn — through the living layers — so the loss falls, nothing crashes, and results are only mysteriously mediocre. Without touching the soil, you would never suspect that a third of your model has been dead weight since the first step.

The cause is usually one accidental line. A connection cut in the wrong place, a layer wired up but never used, or a frozen setting left over from an experiment.

How it works

loss
  │ backward
  ▼
layer 3   soil wet   ✓ gradient arrived
layer 2   soil wet   ✓
──── blockage ────
layer 1   soil DRY   ✗ no gradient — this layer never learns

The check runs after one backward pass and reads one number per layer. Dry soil anywhere is a bug with an address.

Where you have seen this

An electrician hunting a dead socket does not rewire the house. She walks the circuit with a tester pen, outlet by outlet, until the live wire stops being live. Same method, same economy: test each point, find the boundary between working and dead.

Remember this

  • Gradients flow backward from the loss; a cut starves everything behind it.
  • Starved layers do not crash — they quietly stop learning.
  • One loop after backward() reads the soil moisture of every layer.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Captured with torch 2.5.1 on CPU; seeded and deterministic.

The dry-soil detector

This model contains a planted bug — a detach in the middle of the forward pass:

grad_flow.py
import torch
import torch.nn as nn

class Net(nn.Module):
    def __init__(self):
        super().__init__()
        self.first = nn.Linear(10, 32)
        self.middle = nn.Linear(32, 32)
        self.last = nn.Linear(32, 2)

    def forward(self, x):
        h = torch.relu(self.first(x))
        h = torch.relu(self.middle(h)).detach()   # the bug: detach cuts the wire here
        return self.last(h)

torch.manual_seed(0)
model = Net()
loss = nn.CrossEntropyLoss()(model(torch.randn(16, 10)), torch.randint(0, 2, (16,)))
loss.backward()

for name, p in model.named_parameters():
    if p.grad is None:
        print(f"{name:15s} grad: MISSING")
    else:
        print(f"{name:15s} grad norm: {p.grad.norm().item():.6f}")
Output
first.weight    grad: MISSING
first.bias      grad: MISSING
middle.weight   grad: MISSING
middle.bias     grad: MISSING
last.weight     grad norm: 0.184076
last.bias       grad norm: 0.079574

The map reads like the pipe diagram: everything behind the detach is dry. Note the direction — gradients flow backward, so a cut after middle starves middle and first, while last (between the cut and the loss) stays healthy.

Remove the .detach() and rerun:

Output
first.weight    grad norm: 0.083987
first.bias      grad norm: 0.019576
middle.weight   grad norm: 0.142110
middle.bias     grad norm: 0.038178
last.weight     grad norm: 0.184076
last.bias       grad norm: 0.079574

All soil wet. Notice last's numbers are identical in both runs — its water never depended on the broken joint.

Reading the three verdicts

MISSING (grad is None) — no gradient ever arrived. The wire is cut: a detach() or .data in the path, requires_grad=False on those parameters, an operation autograd cannot pass through (argmax, round, comparisons, integer casts, a trip through NumPy), or the layer genuinely is not in the path — defined in __init__ but never called in forward.

Norm exactly 0.0 — a gradient arrived and it is zero. Different disease: dead ReLUs (every input negative), multiplication by an exact zero mask, or a loss term weighted to nothing.

Norms present but vanishingly small in early layers — each layer passes a fraction of the signal and the fractions compound; ten layers of halving leaves a thousandth. That is the vanishing gradient problem — see backpropagation — treated with normalisation, residual connections, or gentler depth.

Common mistakes

Running the check before any backward(). Every grad is None before the first backward pass, and the printout screams catastrophe. The check only means something after a real loss has flowed.

Forgetting grads accumulate. Checking after several backwards without zero_grad reads summed history, not this step's flow. One clean zero_grad, one forward, one backward, then read.

Confusing intentional freezes with bugs. Frozen backbone layers in transfer learning should read MISSING — the check reports flow, and you supply the judgement about which dryness is deliberate.

Using .data to dodge an autograd error. It silences the error by cutting the graph — manufacturing exactly the bug this lesson hunts. The honest alternatives are covered in detach vs no_grad.

Try it yourself

Break the model three new ways, predicting the gradient map before each run: set self.first.requires_grad_(False); replace the middle activation with (h > 0).float(); add a fourth layer to __init__ without calling it in forward. Then check what named_parameters shows for the fourth case and note which verdict it produces.

What to learn next

Researcher — Mathematics and papers.

The compounding view

For a depth-$L$ composition, the gradient at layer $l$ carries the product of downstream Jacobians:

$$\frac{\partial \mathcal{L}}{\partial \theta_l} = \frac{\partial \mathcal{L}}{\partial h_L} \left( \prod_{k=l+1}^{L} J_k \right) \frac{\partial h_l}{\partial \theta_l}, \qquad J_k = \frac{\partial h_k}{\partial h_{k-1}}$$

  • $h_k$ — layer $k$'s activations; $J_k$ — its Jacobian; $\theta_l$ — layer $l$'s parameters.

Singular values of the product compound geometrically: uniformly below 1 vanishes, above 1 explodes (the previous lesson's territory). Glorot and Bengio (2010) derived initialisation scales that keep per-layer variance neutral; He et al. (2015) adjusted for ReLU's halving; residual connections (He et al., 2016, Deep Residual Learning) add an identity term to each $J_k$, guaranteeing an unattenuated gradient path — the architectural fix that made 100+ layers routine. LSTMs' gating (Hochreiter and Schmidhuber, 1997) is the recurrent version of the same idea.

Instruments beyond the one-shot loop

  • Backward hooks — register_full_backward_hook on modules streams per-layer gradient statistics during training rather than post-hoc; the mechanics live in forward and backward hooks.
  • Update-to-weight ratios — $|\eta\, g_l| / |\theta_l|$ per layer, healthy near $10^{-3}$; a profile across layers localises pathology better than raw norms, which scale with layer size.
  • Gradient histograms over time — TensorBoard/W&B log these cheaply; the diagnostic signature of a dying layer is a distribution collapsing toward a spike at zero across epochs.
  • For verifying a custom backward implementation rather than flow, finite-difference checking is the tool — gradcheck.

None vs zero, precisely

p.grad is None after backward means autograd never wrote to that leaf: the parameter was not reached in the graph walk (severed path, requires_grad=False, or unused module). A written zero means the chain rule genuinely evaluated to zero — reachable but locally flat. The distinction matters operationally: None never becomes nonzero by training longer, while zero can (a dead ReLU can revive if other parameters shift its pre-activations). Since PyTorch 2.x, optimizer.zero_grad(set_to_none=True) is the default and resets grads to None for memory savings — so the audit must run after a backward, not after a zeroing, to avoid reading the reset as blockage.

What to learn next