requires_grad and leaf tensors
Only leaf tensors with requires_grad=True keep a .grad after backward — everything in between is used for the walk and thrown away.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A leaf tensor is one you created yourself, and only leaves that asked for gradients keep them after the backward walk.
Think of a family tree. The eldest ancestors at the top were not born from anyone on the tree — everyone else descends from them. Tensors form the same picture. The ones you create directly are the ancestors; every tensor made by a calculation is a descendant.
PyTorch's word for an ancestor is a leaf — a tensor that no recorded operation produced.
Why it exists
Training only ever adjusts the tensors you start with: the weights. The in-between results — this batch's predictions, this batch's hidden values — are scaffolding. They exist for a moment and are rebuilt fresh next batch.
So autograd splits tensors into two roles. Leaves that asked for gradients get their blame stored, so the optimiser can read it and adjust them. Everything downstream has its blame passed through and immediately discarded. Storing blame for millions of temporary values would waste memory on numbers nobody reads.
How it works
w (you made it, asked for grad) -> leaf -> .grad is KEPT
data (you made it, no grad asked) -> leaf -> ignored by backward
h = w * 2 (made by an operation) -> not leaf -> blame passes through
loss = ... (made by an operation) -> not leaf -> blame passes throughAfter the backward walk, exactly one kind of tensor has anything in .grad: a leaf that asked.
A real example you have seen
In a school's results system, the marks flow upward: question marks make subject totals, subject totals make the report card. Nobody files the in-between subtotals — they are recomputed from the question marks any time. The question marks are the leaves: the stored, adjustable source of everything else.
Remember this
- A leaf is a tensor you created, not one an operation produced.
- Only leaves with
requires_grad=Truekeep a.grad. - In-between tensors pass blame along and store nothing.
What to learn next
- detach, no_grad and inference_mode — cutting tensors out of the tree on purpose.
- Gradient accumulation — exploiting the
+=behaviour deliberately. - How the autograd graph is built and freed — the recording these rules govern.
Developer — Code and libraries.
Setup
pip install torchOutputs verified with torch 2.5.1, CPU.
The three kinds of tensor, live
import torch
w = torch.tensor([1., 2., 3.], requires_grad=True) # a leaf: you made it
data = torch.tensor([10., 10., 10.]) # a leaf too, no grad wanted
h = w * 2 # made BY an operation: not a leaf
loss = (h * data).sum()
print(w.is_leaf, data.is_leaf, h.is_leaf)
print(w.requires_grad, data.requires_grad, h.requires_grad)
loss.backward()
print(w.grad) # populated: leaf with requires_grad=True
print(h.grad) # None: non-leaf grads are used, then thrown awayTrue True False True False True tensor([20., 20., 20.]) None
Running this also prints a UserWarning on the last line: "The .grad attribute of a Tensor that is not a leaf Tensor is being accessed. Its .grad attribute won't be populated during autograd.backward()…" — PyTorch telling you, precisely, what this lesson teaches.
The walkthrough
h.requires_grad is True, yet h.grad is None. These are different questions. requires_grad=True on h means "blame flows through me" — it must, or w could never be reached. Keeping the blame is a separate privilege reserved for leaves. During the walk, h's gradient exists momentarily, is used to compute w's, and is freed.
data took part in the maths but wanted no gradient. Inputs and labels stay requires_grad=False — you do not plan to adjust the data to please the model, though that sentence describes a real technique (researcher block).
When you genuinely want an intermediate's gradient — debugging, inspecting attention — ask before the walk:
import torch
w = torch.tensor([1., 2., 3.], requires_grad=True)
h = w * 2
h.retain_grad() # ask autograd to keep this one
(h * 10).sum().backward()
print(h.grad)tensor([10., 10., 10.])
And you cannot flip the switch on a non-leaf:
h = w * 2
h.requires_grad_(False)RuntimeError: you can only change requires_grad flags of leaf variables. If you want to use a computed variable in a subgraph that doesn't require differentiation use var_no_grad = var.detach().
The error's suggestion — detach() — is the subject of the next lesson.
Common mistakes
Reading None from model.layer.weight.grad before any backward. .grad starts as None and first fills during a backward pass. None before training is healthy; None after a backward means blame never reached that leaf — a genuinely broken graph, usually a detach or a re-wrapped tensor somewhere.
Re-wrapping tensors and severing the tree. w2 = torch.tensor(h) creates a new leaf that remembers nothing about w. Blame stops there. Wrapping an existing tensor in torch.tensor(...) mid-computation is almost always a bug (and warns).
Updating weights without no_grad and creating accidental non-leaves. Writing w = w - lr * w.grad makes the new w an operation result — a non-leaf whose .grad will stay None forever after. Optimisers update in-place under no_grad for exactly this reason.
Expecting .grad on every tensor "because requires_grad is True". The two-condition rule — leaf AND requires_grad — is the answer to a whole family of confused forum posts. Print t.is_leaf, t.requires_grad and the mystery usually ends.
Try it yourself
Build a two-step chain a -> b -> c with a a tracked leaf, call backward() on c.sum(), and print all three .grads. Then add b.retain_grad() and verify b.grad appears while c.grad stays None. Work out why the chain rule makes b.grad what it is.
What to learn next
- detach, no_grad and inference_mode — cutting tensors out of the tree on purpose.
- Gradient accumulation — exploiting the
+=behaviour deliberately. - How the autograd graph is built and freed — the recording these rules govern.
Researcher — Mathematics and papers.
Formal definition and the AccumulateGrad node
A leaf is a tensor with grad_fn is None — no producing node in the graph. For each leaf with requires_grad=True, the engine attaches an AccumulateGrad node as the terminal consumer of its adjoint; that node performs leaf.grad += adjoint (allocating on first write). The accumulation — +=, never = — implements the fan-in sum of the chain rule across multiple uses and multiple backward calls, and is what zero_grad() resets. Since torch 1.x, zero_grad(set_to_none=True) is the default: resetting to None skips a memory write and lets the next accumulation allocate lazily.
Non-leaf adjoints are buffers of the engine, consumed by consumer nodes and freed as reference counts drop; retain_grad() registers a hook copying the buffer into .grad before release. tensor.register_hook(fn) generalises this: fn observes (and may replace) the adjoint in flight — the standard tool for gradient inspection and surgery.
requires_grad propagation
requires_grad is an OR over operation inputs: any tracked input makes the output tracked. The flag therefore defines a frontier: the recorded subgraph is exactly the set of ops reachable from tracked leaves. Freezing a submodule via p.requires_grad_(False) per parameter prunes both the tape construction and the backward walk below those nodes — a real compute/memory saving, distinct from (and complementary to) omitting parameters from the optimiser. See the transfer-learning pattern in PyTorch basics.
When inputs do require gradients
Setting requires_grad=True on data is a legitimate family of techniques: adversarial example generation (FGSM — Goodfellow et al., 2015, Explaining and Harnessing Adversarial Examples: x_adv = x + ε·sign(∇_x L)), saliency maps (Simonyan et al., 2014), input optimisation in style transfer (Gatys et al., 2016), and prompt/soft-embedding tuning. The machinery is identical; only which leaves ask differs.
Parameters are leaves by construction
nn.Parameter is a Tensor subclass with requires_grad=True by default, auto-registered by nn.Module attribute assignment. An optimiser is, mechanically, an object holding references to leaves and applying in-place updates under torch.no_grad() from their .grad fields — the entire training loop contract reduces to the leaf rule of this lesson.
References
- Paszke et al. (2017), Automatic differentiation in PyTorch, NIPS Autodiff Workshop — the leaf/AccumulateGrad design.
- PyTorch docs, Autograd mechanics — the authoritative statement of the leaf rules and hook ordering.
- Goodfellow et al. (2015), arXiv:1412.6572 — gradients with respect to inputs, weaponised.
What to learn next
- detach, no_grad and inference_mode — cutting tensors out of the tree on purpose.
- Gradient accumulation — exploiting the
+=behaviour deliberately. - How the autograd graph is built and freed — the recording these rules govern.