AI glossary

Vanishing gradient

In one sentence The vanishing gradient problem is when the training signal fades to nothing as it travels back through many layers, so early layers stop learning.

By Updated

The vanishing gradient problem is when the error signal shrinks layer by layer as it flows backwards, until the early layers receive almost nothing and stop learning.

Play whisper-down-the-lane with thirty children. The first child hears the full sentence. By the tenth, half the words are mangled. By the thirtieth, nothing meaningful survives. Backpropagation passes the error signal backwards through the layers the same way, and at each layer the signal gets multiplied by local factors. If those factors are mostly below one, the signal shrinks exponentially with depth.

The result is a network whose last layers train happily while its first layers barely move — and the first layers are the ones that read the raw input. For years this capped how deep networks could usefully be.

Three inventions largely tamed it. The ReLU activation-function replaced sigmoids, whose slopes are at most 0.25 and crush the signal at every layer. Residual-connections give the gradient a shortcut path that skips the multiplications entirely. And normalisation layers such as batch-normalization keep the signal in a healthy range. In recurrent networks, the LSTM was designed specifically to carry the signal across many time steps.

Its mirror image is the exploding gradient, where the factors are above one and the signal blows up — handled by gradient-clipping.

Where to go next