AI glossary

Backpropagation

In one sentence Backpropagation works out how much each weight in a network contributed to the error, by passing the error backwards through the layers.

By Updated

Backpropagation is the method that works out how much each individual weight contributed to the model's error, by carrying the error backwards through the layers.

A dish comes out too salty. You do not throw away the kitchen. You walk back through the steps — the marinade, the gravy, the final seasoning — and ask at each step how much salt it added, so you know which step to change and by how much next time. Backpropagation does the same walk backwards through a network, and at every weight it answers one question: if this number moved up a little, would the error go up or down, and by how much?

That answer is the gradient. Getting it is only possible because each layer is a simple function applied to the layer before it. That structure lets the chain rule of calculus split the error at the end among the layers behind it, one hop at a time. Each layer receives the blame passed down from above, keeps the share that belongs to its own weights, and passes the rest further back.

Forward, then backward

Forward :  input → layer 1 → layer 2 → prediction → loss
Backward:                                          loss
                 ← layer 1 grads ← layer 2 grads ←
Then    :  gradient descent nudges every weight

Backpropagation only computes the gradients. Changing the weights is a separate job, done by gradient descent and the optimizer. In PyTorch the split is visible: loss.backward() fills in the gradients, optimizer.step() applies them, and optimizer.zero_grad() clears them so the next batch starts clean. Forgetting that last call is one of the most common training bugs there is — gradients accumulate by default, so the model ends up updating on a growing sum of old batches.

Where to go next