element 0 of tensors does not require grad and does not have a grad_fn
You called .backward() on a tensor with no connection to any trainable weight. Usually the loss was computed inside torch.no_grad(), or a step in the middle broke the graph.
Updated
The error
Traceback (most recent call last):
File "train.py", line 41, in <module>
loss.backward()
RuntimeError: element 0 of tensors does not require grad and does not have a grad_fnTwo errors that come from the same misunderstanding:
RuntimeError: Trying to backward through the graph a second time (or directly access saved tensors after they have already been freed).
ValueError: optimizer got an empty parameter list
What it means
To compute gradients, PyTorch records every operation performed on tensors that need them, building a graph as your forward pass runs. loss.backward() walks that graph backwards.
Your loss tensor has no graph attached. grad_fn is None, which means PyTorch has no record of how this number was produced, so there is nothing to walk backwards through and no weight to assign blame to. The number is a plain number, not a result.
You can confirm it in one line before the failing call:
print(loss.requires_grad, loss.grad_fn)False None
Why it happens
The loss was computed inside torch.no_grad(). This is by far the most common cause. The context manager switches off graph recording, so everything computed inside it comes out detached. It usually happens when a validation block's indentation swallows the training step, or when someone adds no_grad to save memory without noticing it now wraps the loss.
Something in the chain detached the tensor. .item(), .detach(), .numpy(), int(), float() and torch.tensor(existing_tensor) all produce a value with no history. A loss built from preds.detach() or from loss_value = loss.item() cannot be backpropagated.
Every parameter is frozen. If all parameters have requires_grad = False — after freezing a backbone and forgetting to unfreeze the new head — nothing in the model requires gradients, so nothing downstream does either.
A non-differentiable operation sits in the middle. argmax, round, comparisons, and casting to an integer type all break the chain. Converting logits to a predicted class and computing accuracy gives a number you cannot train on, because there is no smooth relationship between the weights and that number.
torch.inference_mode() was used instead of no_grad(). Tensors created inside it are marked more strictly, and using them in a training computation later fails even outside the block.
How to fix it
1. Check what is wrapping the loss computation. Look at the indentation around the failing line.
# wrong — the loss is built with graph recording switched off
with torch.no_grad():
preds = model(x)
loss = criterion(preds, y)
loss.backward() # this error
# correct — no_grad belongs around evaluation only
preds = model(x)
loss = criterion(preds, y)
loss.backward()
optimizer.step()
optimizer.zero_grad(set_to_none=True)A validation loop that shares a function with training is a frequent source: check that a no_grad added for evaluation is not being applied to both.
2. Trace where the graph was lost. Print as you go and find the first False.
preds = model(x)
print("preds:", preds.requires_grad, preds.grad_fn)
loss = criterion(preds, y)
print("loss :", loss.requires_grad, loss.grad_fn)The first line that prints False None is the operation that cut the connection.
3. Remove the detaching call. Common repairs:
loss = criterion(preds, y) # not criterion(preds.detach(), y)
total = total + loss # for building a combined loss
running = running + loss.item() # .item() is correct for LOGGING onlyKeep the two apart: the tensor goes into backward(), the float goes into your log. If you need a copy of a tensor while keeping the graph, use x.clone() — torch.tensor(x) copies the values and drops the history.
4. Confirm something in the model is actually trainable.
trainable = [n for n, p in model.named_parameters() if p.requires_grad]
print(f"{len(trainable)} trainable tensors")
print(trainable[:5])If that prints zero, unfreeze what you meant to train:
for p in model.parameters():
p.requires_grad = False # freeze the backbone
for p in model.fc.parameters():
p.requires_grad = True # train the new headAnd build the optimizer from the parameters that are actually training, which also fixes optimizer got an empty parameter list:
optimizer = torch.optim.Adam(
(p for p in model.parameters() if p.requires_grad), lr=1e-3
)5. Do not train through argmax. Accuracy is not a loss.
# wrong — argmax has no usable gradient
preds = logits.argmax(dim=1)
loss = (preds != y).float().mean()
# correct — cross-entropy on the raw scores is differentiable
loss = nn.CrossEntropyLoss()(logits, y)Compute metrics from argmax for reporting, and compute the loss from the raw logits. If you genuinely need a discrete decision inside the model, that calls for a different technique such as Gumbel-softmax, not a small patch here.
6. When using gradient checkpointing with a frozen base model, enable input gradients. This is the specific cause behind this error in LoRA and PEFT fine-tuning, and the fix is a documented one-liner:
model.gradient_checkpointing_enable()
model.enable_input_require_grads() # without this, checkpointed blocks have no grad pathCheckpointing recomputes activations during the backward pass, and it needs at least one input to the block to require gradients in order to reconnect the graph. With every base weight frozen, that condition is otherwise unmet.
7. If the message is about backward being called twice, you are reusing a graph that was already freed. Either the loss is accumulated across iterations without detaching, or the same loss is passed to backward() more than once. Recompute the forward pass each iteration, and detach anything carried between steps — hidden states in an RNN are the classic case:
hidden = tuple(h.detach() for h in hidden) # carry the values, not the historyHow to prevent it
Keep training and evaluation in separate functions, with torch.no_grad() appearing only in the evaluation one. Mixing them in one function with a flag is what puts a context manager in the wrong place.
Use the decorator form for evaluation functions, where the scope is unmistakable:
@torch.no_grad()
def evaluate(model, loader):
model.eval()
...Adopt one habit for logging: anything you store, print or append gets .item() or .detach().cpu(), and anything you optimise stays untouched. That single rule prevents this error and also prevents the memory leak in CUDA out of memory, which comes from the opposite mistake.
Finally, run one training step on two samples before any long run and assert the loss has a grad_fn. Four seconds of checking beats finding out after the data loader has warmed up.
Related errors
- Expected all tensors to be on the same device
- Loss is NaN during training
- CUDA out of memory — caused by keeping the graph when you meant to drop it, the mirror image of this error