Autograd in Depth

In-place operations and when they break autograd

Operations ending in an underscore overwrite their tensor to save memory — and when they overwrite a value autograd saved for backward, the walk fails with a version error.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

An in-place operation overwrites a tensor's numbers where they sit — saving memory, and sometimes destroying evidence autograd was counting on.

Think of solving a maths exam in the margin of the question paper, rubbing out each rough step to make room for the next. You save paper. But when the teacher asks you to show your working, the working is gone — you erased the very steps the checking needs.

Autograd is that teacher. During the forward pass it asks you to leave certain values untouched, because the backward walk will need to re-read them.

Why it exists

Normal operations make a new tensor for every result — new memory each time. In-place versions overwrite the input instead. In PyTorch they are marked with an underscore: add_, relu_, clamp_.

For huge tensors that thrift is real. But the backward walk replays the forward recording, and the recording often includes pointers to values, not copies of them. Overwrite a pointed-at value, and the replay would silently compute wrong blame — so PyTorch keeps a tamper-seal on every tensor and refuses to replay from edited evidence.

How it works

forward:   a ──(exp)──> b        autograd notes: "keep b, I need it later"
edit:      b.add_(1)             the kept value is overwritten  [seal broken]
backward:  replay reaches b  ->  seal check fails  ->  loud error

The seal is a counter on each tensor, ticking up on every in-place edit. The recording remembers which count it expects; a mismatch stops everything.

A real example you have seen

Bank statements work on the same principle. The bank never edits an entry in place — corrections are new entries, because auditors must be able to replay history. A ledger you overwrite is a ledger you cannot audit. Autograd audits your forward pass, and demands the same discipline.

Remember this

  • Underscore operations (add_, relu_) overwrite instead of creating.
  • Autograd saves values for the backward walk; overwriting one breaks the walk.
  • The error is loud and names the problem — silent wrong gradients are what it prevents.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch

Outputs verified with torch 2.5.1, CPU.

The three cases that cover everything

inplace.py
import torch

# Case 1: in-place on a leaf that requires grad - refused instantly
w = torch.tensor([1., 2.], requires_grad=True)
try:
    w.add_(1)
except RuntimeError as err:
    print("case 1:", err)

# Case 2: in-place on a value autograd saved - caught at backward time
a = torch.tensor([1., 2.], requires_grad=True)
b = a.exp()          # exp saves its OUTPUT, because d(exp)/da = exp(a)
b.add_(1)            # overwrites that saved value
try:
    b.sum().backward()
except RuntimeError as err:
    print("case 2:", str(err)[:130], "...")

# Case 3: in-place that touches nothing saved - works fine
p = torch.tensor([1., 2.], requires_grad=True)
q = p * 5            # multiply-by-constant saves no tensor
q.add_(10)
q.sum().backward()
print("case 3:", p.grad)
Output
case 1: a leaf Variable that requires grad is being used in an in-place operation.
case 2: one of the variables needed for gradient computation has been modified by an inplace operation: [torch.FloatTensor [2]], which is  ...
case 3: tensor([5., 5.])

The walkthrough

Case 1 fails immediately. Editing a weight in-place mid-graph would make the recorded history a lie, so leaves that require grad reject underscore ops outright. (Optimisers edit weights in-place legally by doing it under no_grad, outside the recording — see detach and no_grad.)

Case 2 is the famous one. exp is its own derivative, so autograd saves the output b for the backward walk. b.add_(1) overwrites it. Nothing complains yet — the crime is detected at replay time. The full error even tells you where you are: "…which is output 0 of ExpBackward0, is at version 1; expected version 0 instead", and suggests torch.autograd.set_detect_anomaly(True), which re-runs with tracking and points at the guilty forward line.

Case 3 shows the rule is precise, not paranoid. Multiplying by a constant needs no saved tensor — the derivative is the constant. Editing q afterwards destroys no evidence, so the walk succeeds. Whether an in-place op is legal depends on what the surrounding operations saved, which is why the same add_ is fine here and fatal one case up.

The tamper-seal is visible, if you are curious:

python
import torch
a = torch.tensor([1., 2.], requires_grad=True)
b = a.exp()
print(b._version)
b.add_(1)
print(b._version)   # every in-place edit bumps the counter
Output
0
1

Common mistakes

"Fixing" the error with .data. Old advice suggests b.data.add_(1) — it edits the values while dodging the version counter. The error disappears and the gradients become silently wrong: the exact disaster the seal exists to prevent. Never trade a loud error for a quiet lie.

In-place on views. A view shares memory with its base, so x[0].add_(1) or editing a slice edits the base tensor — and can break a save you made three lines earlier through a different name. When a version error points somewhere confusing, look for aliases of the tensor, not only the tensor.

Sprinkling relu_(inplace=True) everywhere for speed. nn.ReLU(inplace=True) is legal only when the input is not saved by anything else — residual connections and reused activations are where it blows up. Measure first: activation memory savings are real in tight memory budgets, but the debugging cost of a misplaced one is realer.

Accumulating with += on a tracked tensor. total += loss on a tensor that requires grad is an in-place add participating in the graph — legal but graph-growing, and occasionally version-error-adjacent. For running metrics, total += loss.item().

Try it yourself

Change case 2's exp() to * 5 and predict what happens before running. Then restore exp() and fix case 2 properly — with an out-of-place b = b + 1 — and confirm the gradient is what the chain rule says it should be for exp(a) + 1.

What to learn next

Researcher — Mathematics and papers.

The version-counter protocol

Every tensor's storage carries _version, incremented by any mutating op. When a node saves a tensor for backward, it records (pointer, version-at-save). On unpacking during the backward walk, current version ≠ recorded version raises the case-2 error. The check is conservative and sound: it flags every mutation of saved storage, including mutations that happen not to change the bytes the backward reads (aliased views), preferring false alarms to wrong gradients. inference_mode skips maintaining these counters entirely — the source of both its speed and its export restrictions (detach lesson).

Which tensors an op saves is a per-op derivative-formula fact, catalogued in derivatives.yaml in the PyTorch source: mul(a, b) saves both operands (each one's gradient needs the other); exp saves its output (ȳ·y); add saves nothing (adjoint passes through). Case-by-case legality of in-place follows mechanically from that table — the developer block's three cases are its first three rows in disguise.

In-place on views: grad_fn rewriting

Mutating a view is mutating the base, so autograd rewrites history: after v = x[0]; v.add_(1), the graph replaces the view's node using AsStridedBackward-style machinery, and the base's version bumps for all aliases. This is why alias analysis, not local inspection, is needed to reason about a version error — and why the error message reports the tensor's type and shape rather than a variable name: several names may share the storage.

Memory economics

The saving from in-place ops is bounded by activation memory, which for deep nets dominates parameters (O(Σ layer outputs) — see graph mechanics). But allocator behaviour blunts naive savings: PyTorch's caching allocator reuses freed blocks efficiently, so out-of-place code often costs less than intuition suggests. The systematic memory lever is checkpointing (Chen et al., 2016, arXiv:1604.06174), not underscore-hunting. Empirically, in-place matters most for (a) very large element-wise chains and (b) frameworks-within-frameworks that forbid allocation, which is why compiler passes handle it below your code.

Functionalization: making in-place disappear

torch.compile and the export stack run functionalization: a pass rewriting every in-place op into an out-of-place op plus explicit aliasing bookkeeping, yielding a pure dataflow graph for optimisation, then re-fusing where profitable. Consequence worth knowing: under the compiler, hand-written in-place "optimisations" are frequently no-ops — the pass undoes and redoes such decisions with a global view. The eager-mode rules in this lesson remain exactly true; the compiled world quietly makes many of them moot.

References

  • PyTorch docs, Autograd mechanics, §"In-place operations with autograd" and §"In-place correctness checks" — the normative protocol.
  • Chen et al. (2016), Training Deep Nets with Sublinear Memory Cost — the principled alternative to in-place thrift.
  • Bernstein et al., PyTorch RFC: Functionalization (pytorch/rfcs) — design notes for the compiler-era treatment.

What to learn next