Error database

element 0 of tensors does not require grad and does not have a grad_fn

You called .backward() on a tensor with no connection to any trainable weight. Usually the loss was computed inside torch.no_grad(), or a step in the middle broke the graph.

The message you saw
element 0 of tensors does not require grad and does not have a grad_fn

By Updated

The error

Output
Traceback (most recent call last):
  File "train.py", line 41, in <module>
    loss.backward()
RuntimeError: element 0 of tensors does not require grad and does not have a grad_fn

Two errors that come from the same misunderstanding:

Output
RuntimeError: Trying to backward through the graph a second time (or directly access saved tensors after they have already been freed).
Output
ValueError: optimizer got an empty parameter list

What it means

To compute gradients, PyTorch records every operation performed on tensors that need them, building a graph as your forward pass runs. loss.backward() walks that graph backwards.

Your loss tensor has no graph attached. grad_fn is None, which means PyTorch has no record of how this number was produced, so there is nothing to walk backwards through and no weight to assign blame to. The number is a plain number, not a result.

You can confirm it in one line before the failing call:

python
print(loss.requires_grad, loss.grad_fn)
Output
False None

Why it happens

The loss was computed inside torch.no_grad(). This is by far the most common cause. The context manager switches off graph recording, so everything computed inside it comes out detached. It usually happens when a validation block's indentation swallows the training step, or when someone adds no_grad to save memory without noticing it now wraps the loss.

Something in the chain detached the tensor. .item(), .detach(), .numpy(), int(), float() and torch.tensor(existing_tensor) all produce a value with no history. A loss built from preds.detach() or from loss_value = loss.item() cannot be backpropagated.

Every parameter is frozen. If all parameters have requires_grad = False — after freezing a backbone and forgetting to unfreeze the new head — nothing in the model requires gradients, so nothing downstream does either.

A non-differentiable operation sits in the middle. argmax, round, comparisons, and casting to an integer type all break the chain. Converting logits to a predicted class and computing accuracy gives a number you cannot train on, because there is no smooth relationship between the weights and that number.

torch.inference_mode() was used instead of no_grad(). Tensors created inside it are marked more strictly, and using them in a training computation later fails even outside the block.

How to fix it

1. Check what is wrapping the loss computation. Look at the indentation around the failing line.

python
# wrong — the loss is built with graph recording switched off
with torch.no_grad():
    preds = model(x)
    loss = criterion(preds, y)
loss.backward()                       # this error

# correct — no_grad belongs around evaluation only
preds = model(x)
loss = criterion(preds, y)
loss.backward()
optimizer.step()
optimizer.zero_grad(set_to_none=True)

A validation loop that shares a function with training is a frequent source: check that a no_grad added for evaluation is not being applied to both.

2. Trace where the graph was lost. Print as you go and find the first False.

python
preds = model(x)
print("preds:", preds.requires_grad, preds.grad_fn)

loss = criterion(preds, y)
print("loss :", loss.requires_grad, loss.grad_fn)

The first line that prints False None is the operation that cut the connection.

3. Remove the detaching call. Common repairs:

python
loss = criterion(preds, y)                 # not criterion(preds.detach(), y)
total = total + loss                       # for building a combined loss
running = running + loss.item()            # .item() is correct for LOGGING only

Keep the two apart: the tensor goes into backward(), the float goes into your log. If you need a copy of a tensor while keeping the graph, use x.clone() — torch.tensor(x) copies the values and drops the history.

4. Confirm something in the model is actually trainable.

python
trainable = [n for n, p in model.named_parameters() if p.requires_grad]
print(f"{len(trainable)} trainable tensors")
print(trainable[:5])

If that prints zero, unfreeze what you meant to train:

python
for p in model.parameters():
    p.requires_grad = False               # freeze the backbone
for p in model.fc.parameters():
    p.requires_grad = True                # train the new head

And build the optimizer from the parameters that are actually training, which also fixes optimizer got an empty parameter list:

python
optimizer = torch.optim.Adam(
    (p for p in model.parameters() if p.requires_grad), lr=1e-3
)

5. Do not train through argmax. Accuracy is not a loss.

python
# wrong — argmax has no usable gradient
preds = logits.argmax(dim=1)
loss = (preds != y).float().mean()

# correct — cross-entropy on the raw scores is differentiable
loss = nn.CrossEntropyLoss()(logits, y)

Compute metrics from argmax for reporting, and compute the loss from the raw logits. If you genuinely need a discrete decision inside the model, that calls for a different technique such as Gumbel-softmax, not a small patch here.

6. When using gradient checkpointing with a frozen base model, enable input gradients. This is the specific cause behind this error in LoRA and PEFT fine-tuning, and the fix is a documented one-liner:

python
model.gradient_checkpointing_enable()
model.enable_input_require_grads()        # without this, checkpointed blocks have no grad path

Checkpointing recomputes activations during the backward pass, and it needs at least one input to the block to require gradients in order to reconnect the graph. With every base weight frozen, that condition is otherwise unmet.

7. If the message is about backward being called twice, you are reusing a graph that was already freed. Either the loss is accumulated across iterations without detaching, or the same loss is passed to backward() more than once. Recompute the forward pass each iteration, and detach anything carried between steps — hidden states in an RNN are the classic case:

python
hidden = tuple(h.detach() for h in hidden)     # carry the values, not the history

How to prevent it

Keep training and evaluation in separate functions, with torch.no_grad() appearing only in the evaluation one. Mixing them in one function with a flag is what puts a context manager in the wrong place.

Use the decorator form for evaluation functions, where the scope is unmistakable:

python
@torch.no_grad()
def evaluate(model, loader):
    model.eval()
    ...

Adopt one habit for logging: anything you store, print or append gets .item() or .detach().cpu(), and anything you optimise stays untouched. That single rule prevents this error and also prevents the memory leak in CUDA out of memory, which comes from the opposite mistake.

Finally, run one training step on two samples before any long run and assert the loss has a grad_fn. Four seconds of checking beats finding out after the data loader has warmed up.