RuntimeError: one of the variables needed for gradient computation has been modified by an inplace operation
You edited a tensor in place after autograd saved it for the backward pass. Turn on anomaly detection to find the line, then replace the in-place edit with an out-of-place one.
Updated
The error
RuntimeError: one of the variables needed for gradient computation has been modified by an inplace operation: [torch.FloatTensor [64, 128]], which is output 0 of ReluBackward0, is at version 2; expected version 1 instead. Hint: enable anomaly detection to find the operation that failed to compute its gradient, with torch.autograd.set_detect_anomaly(True).
What it means
To compute gradients, autograd saves certain intermediate tensors during the forward pass. An in-place operation — anything that rewrites a tensor's memory, like x += 1, x[mask] = 0 or methods ending in _ — changed one of those saved tensors before backward ran. The saved value is now wrong, so the gradient would be wrong. PyTorch tracks a version counter per tensor to catch exactly this, and stops.
The message even tells you which tensor: here, the output of a ReLU, edited once after being saved.
Why it happens
In-place edits look harmless and often are — until the edited tensor happens to be one autograd needed. Frequent culprits:
x += somethingon an activation insideforwardnn.ReLU(inplace=True)colliding with a layer that needs its input preserved- Masked assignment:
out[out < 0] = 0 - Normalising in place:
x /= x.norm() - Editing model outputs before the loss:
preds[:, 0] = 0
How to fix it
1. Find the exact line first. The error surfaces at backward(), far from the cause. Anomaly detection makes the forward line appear in the traceback:
torch.autograd.set_detect_anomaly(True) # temporarily, at the topRun once, read the second traceback it prints, then remove it — it slows training a lot.
2. Replace the in-place edit with an out-of-place version.
x = x + residual # was: x += residual
out = torch.clamp(out, min=0) # was: out[out < 0] = 0
out = torch.where(mask, torch.zeros_like(out), out) # was: out[mask] = 0
x = x / x.norm() # was: x /= x.norm()Same maths, new tensor, saved values untouched.
3. Switch off inplace=True on activations if the traceback points there.
self.act = nn.ReLU() # was: nn.ReLU(inplace=True)The memory saving is minor in most models; correctness first.
4. To edit values for logging or visualisation, work on a copy.
display = preds.detach().clone()
display[:, 0] = 0.detach().clone() gives an independent tensor outside the graph, free to mutate.
How to prevent it
Inside anything autograd will traverse — forward methods and loss computation — default to out-of-place operations. Reserve in-place tricks for torch.no_grad() blocks (like manual weight updates), where autograd is not watching and they are safe.