RuntimeError: grad can be implicitly created only for scalar outputs
You called backward() on a tensor with more than one element. Reduce the loss to a single number first with .mean() or .sum() — usually a reduction='none' left switched on.
Updated
The error
RuntimeError: grad can be implicitly created only for scalar outputs
What it means
loss.backward() computes gradients of the loss with respect to every parameter. That question only has a direct answer when the loss is a single number. Your loss is a vector — one value per sample — and PyTorch cannot know how to weight the elements against each other. It could assume equal weights, but guessing here hides real modelling decisions, so it refuses.
Why it happens
The usual trigger is a loss function created with reduction="none":
criterion = nn.CrossEntropyLoss(reduction="none")
loss = criterion(logits, targets) # shape: [batch_size]
loss.backward() # boomreduction="none" exists for per-sample weighting schemes. It gets copied from such code into ordinary training loops, where it breaks backward.
Hand-rolled losses cause the same thing: computing (pred - target) ** 2 gives a per-element tensor, and forgetting the final mean leaves it unreduced.
How to fix it
1. Reduce to a scalar before backward.
loss = criterion(logits, targets).mean()
loss.backward().mean() is the standard choice — it makes the loss magnitude independent of batch size. .sum() also works but couples gradient scale to batch size, which interacts with the learning rate.
2. Or drop reduction="none" if you never needed it.
criterion = nn.CrossEntropyLoss() # default reduction='mean'3. Keeping per-sample losses on purpose? Weight, then reduce.
per_sample = criterion(logits, targets) # reduction='none'
loss = (per_sample * sample_weights).mean() # explicit weighting
loss.backward()This is the legitimate pattern reduction="none" exists for — the reduction still happens, under your control.
4. The gradient= escape hatch is for vector-Jacobian products. loss.backward(gradient=torch.ones_like(loss)) runs, and equals loss.sum().backward(). If you reach for it to silence this error, write the .sum() or .mean() instead — it says what you mean.
How to prevent it
Print loss.shape once when writing any new training loop; it must be torch.Size([]) — a 0-dimensional scalar. When borrowing loss code from papers or repositories, check the reduction argument first: it encodes an assumption about the surrounding loop.
Related errors
- element 0 of tensors does not require grad — the other common backward() failure
- Trying to backward through the graph a second time
- Loss is NaN during training