Error database

RuntimeError: grad can be implicitly created only for scalar outputs

You called backward() on a tensor with more than one element. Reduce the loss to a single number first with .mean() or .sum() — usually a reduction='none' left switched on.

The message you saw
RuntimeError: grad can be implicitly created only for scalar outputs

By Updated

The error

Output
RuntimeError: grad can be implicitly created only for scalar outputs

What it means

loss.backward() computes gradients of the loss with respect to every parameter. That question only has a direct answer when the loss is a single number. Your loss is a vector — one value per sample — and PyTorch cannot know how to weight the elements against each other. It could assume equal weights, but guessing here hides real modelling decisions, so it refuses.

Why it happens

The usual trigger is a loss function created with reduction="none":

python
criterion = nn.CrossEntropyLoss(reduction="none")
loss = criterion(logits, targets)      # shape: [batch_size]
loss.backward()                        # boom

reduction="none" exists for per-sample weighting schemes. It gets copied from such code into ordinary training loops, where it breaks backward.

Hand-rolled losses cause the same thing: computing (pred - target) ** 2 gives a per-element tensor, and forgetting the final mean leaves it unreduced.

How to fix it

1. Reduce to a scalar before backward.

python
loss = criterion(logits, targets).mean()
loss.backward()

.mean() is the standard choice — it makes the loss magnitude independent of batch size. .sum() also works but couples gradient scale to batch size, which interacts with the learning rate.

2. Or drop reduction="none" if you never needed it.

python
criterion = nn.CrossEntropyLoss()      # default reduction='mean'

3. Keeping per-sample losses on purpose? Weight, then reduce.

python
per_sample = criterion(logits, targets)          # reduction='none'
loss = (per_sample * sample_weights).mean()      # explicit weighting
loss.backward()

This is the legitimate pattern reduction="none" exists for — the reduction still happens, under your control.

4. The gradient= escape hatch is for vector-Jacobian products. loss.backward(gradient=torch.ones_like(loss)) runs, and equals loss.sum().backward(). If you reach for it to silence this error, write the .sum() or .mean() instead — it says what you mean.

How to prevent it

Print loss.shape once when writing any new training loop; it must be torch.Size([]) — a 0-dimensional scalar. When borrowing loss code from papers or repositories, check the reduction argument first: it encodes an assumption about the surrounding loop.