RuntimeError: Trying to backward through the graph a second time
You called backward twice over the same computation, and the graph was freed after the first call. Usually a hidden state or accumulated loss is carrying old graph across iterations — detach it.
Updated
The error
RuntimeError: Trying to backward through the graph a second time (or directly access saved tensors after they have already been freed). Saved intermediate values of the graph are freed when you call .backward() or autograd.grad(). Specify retain_graph=True if you need to backward through the graph a second time or if you need to access saved tensors after calling backward.
What it means
During the forward pass, PyTorch builds a computation graph — a record of every operation, kept so backpropagation can compute gradients. To save memory, backward() frees that record as it goes. Your code then ran backward through some part of the same record again. The saved values are gone, so PyTorch stops.
The message suggests retain_graph=True. Treat that as a diagnosis, not a fix — in most real cases the second backward is an accident.
Why it happens
Some tensor computed in one iteration is still connected to the graph when the next iteration uses it. Each new backward() then reaches back through the old, already-freed graph. The usual carriers:
- An RNN hidden state passed from batch to batch.
- An accumulated loss tensor:
total = total + losskeeps every batch's graph alive, and a latertotal.backward()re-traverses freed pieces. - GAN training where the discriminator step backwards through the generator's graph, and the generator step tries again.
- A tensor computed once outside the loop but used in the loss every iteration.
How to fix it
1. Detach anything that crosses an iteration boundary.
hidden = model.init_hidden(batch_size)
for x, y in loader:
hidden = hidden.detach() # cut the link to last batch's graph
out, hidden = model(x, hidden)
loss = criterion(out, y)
optimizer.zero_grad(set_to_none=True)
loss.backward()
optimizer.step().detach() keeps the values and drops the graph history. This is the standard pattern called truncated backpropagation through time.
2. Accumulate numbers, not graph.
total_loss += loss.item() # plain float for loggingCall backward() on each batch's own loss, inside the loop.
3. In GANs, detach the generator output for the discriminator step.
fake = generator(z)
d_loss = criterion(discriminator(fake.detach()), fake_labels) # D step: no G graph
d_loss.backward()
...
g_loss = criterion(discriminator(fake), real_labels) # G step: fresh use
g_loss.backward()4. Move loop-invariant computation inside the loop — or detach it. If a tensor genuinely must be recomputed against fresh graph each iteration, compute it each iteration.
5. Use retain_graph=True only when two backward passes over one graph are the actual design. Needing it every iteration of a training loop means a leak like the above; memory use will climb until the job dies.
How to prevent it
At the end of each loop body, ask of every tensor that survives to the next iteration: does this still need its graph? Almost always the answer is no — detach it or store .item(). This habit also prevents the slow memory leak that accompanies accidental graph retention.
Related errors
- element 0 of tensors does not require grad — the opposite problem: no graph at all
- one of the variables has been modified by an inplace operation
- grad can be implicitly created only for scalar outputs
- CUDA out of memory — what accumulated graphs cause first