AI glossary

Gradient descent

In one sentence Gradient descent improves a model by repeatedly taking a small step in whichever direction lowers the error.

By Updated

Gradient descent is the method that trains almost every model: check which direction lowers the error, take a small step that way, repeat.

You are on a hillside in thick fog, trying to reach the valley. You cannot see the bottom. What you can do is feel which way the ground slopes under your feet and step downhill. Do that a few thousand times and you arrive somewhere low, without ever having seen the whole landscape. Gradient descent has exactly this much information: the slope at the point where it currently stands.

The "gradient" is that slope, computed by backpropagation for every weight in the model. The "descent" is subtracting a small fraction of it from each weight.

Step size is the setting that matters

error
  │╲                            ╱
  │ ╲  ●→●→●→●                 ╱     good learning rate: steady progress
  │  ╲__________●____________╱
  │
  │╲     ●───────────────→●   ╱      too large: it leaps past the bottom
  │ ╲                        ╱       and the error can blow up to NaN
  │  ╲______________________╱

That fraction is the learning rate, and it is the hyperparameter people spend the most time on. Too small and training crawls. Too large and each step overshoots the valley, the error climbs instead of falling, and you end up looking at a loss that has gone to NaN. A common starting point is 0.001 with the Adam optimizer, then adjust based on what the loss curve does in the first few hundred steps.

In practice nobody computes the slope over the entire dataset — that would be one step per full pass. Instead the slope is estimated from one batch at a time, which is called stochastic gradient descent. The estimate is noisy, and that noise turns out to help, because it shakes the model out of shallow dips that a perfectly smooth path would settle into. Optimizers like Adam add momentum and a per-weight step size on top of this basic idea.

Where to go next

Learn this properly

Full lessons that use this term in context.

  • Mathematics for AI

    Optimization

    The gradient tells you which way to move. Optimization decides how far to move, and it is the difference between a model that trains and one that never does.

  • Deep Learning

    Gradient descent

    Gradient descent finds good weights by repeatedly taking a small step in the direction that lowers the error fastest.

Back to the glossary