Error database

RuntimeError: one of the variables needed for gradient computation has been modified by an inplace operation

You edited a tensor in place after autograd saved it for the backward pass. Turn on anomaly detection to find the line, then replace the in-place edit with an out-of-place one.

The message you saw
RuntimeError: one of the variables needed for gradient computation has been modified by an inplace operation

By Updated

The error

Output
RuntimeError: one of the variables needed for gradient computation has been modified by an inplace operation: [torch.FloatTensor [64, 128]], which is output 0 of ReluBackward0, is at version 2; expected version 1 instead. Hint: enable anomaly detection to find the operation that failed to compute its gradient, with torch.autograd.set_detect_anomaly(True).

What it means

To compute gradients, autograd saves certain intermediate tensors during the forward pass. An in-place operation — anything that rewrites a tensor's memory, like x += 1, x[mask] = 0 or methods ending in _ — changed one of those saved tensors before backward ran. The saved value is now wrong, so the gradient would be wrong. PyTorch tracks a version counter per tensor to catch exactly this, and stops.

The message even tells you which tensor: here, the output of a ReLU, edited once after being saved.

Why it happens

In-place edits look harmless and often are — until the edited tensor happens to be one autograd needed. Frequent culprits:

  • x += something on an activation inside forward
  • nn.ReLU(inplace=True) colliding with a layer that needs its input preserved
  • Masked assignment: out[out < 0] = 0
  • Normalising in place: x /= x.norm()
  • Editing model outputs before the loss: preds[:, 0] = 0

How to fix it

1. Find the exact line first. The error surfaces at backward(), far from the cause. Anomaly detection makes the forward line appear in the traceback:

python
torch.autograd.set_detect_anomaly(True)     # temporarily, at the top

Run once, read the second traceback it prints, then remove it — it slows training a lot.

2. Replace the in-place edit with an out-of-place version.

python
x = x + residual            # was: x += residual
out = torch.clamp(out, min=0)                   # was: out[out < 0] = 0
out = torch.where(mask, torch.zeros_like(out), out)   # was: out[mask] = 0
x = x / x.norm()            # was: x /= x.norm()

Same maths, new tensor, saved values untouched.

3. Switch off inplace=True on activations if the traceback points there.

python
self.act = nn.ReLU()        # was: nn.ReLU(inplace=True)

The memory saving is minor in most models; correctness first.

4. To edit values for logging or visualisation, work on a copy.

python
display = preds.detach().clone()
display[:, 0] = 0

.detach().clone() gives an independent tensor outside the graph, free to mutate.

How to prevent it

Inside anything autograd will traverse — forward methods and loss computation — default to out-of-place operations. Reserve in-place tricks for torch.no_grad() blocks (like manual weight updates), where autograd is not watching and they are safe.