Gradient descent
In one sentence Gradient descent improves a model by repeatedly taking a small step in whichever direction lowers the error.
Updated
Gradient descent is the method that trains almost every model: check which direction lowers the error, take a small step that way, repeat.
You are on a hillside in thick fog, trying to reach the valley. You cannot see the bottom. What you can do is feel which way the ground slopes under your feet and step downhill. Do that a few thousand times and you arrive somewhere low, without ever having seen the whole landscape. Gradient descent has exactly this much information: the slope at the point where it currently stands.
The "gradient" is that slope, computed by backpropagation for every weight in the model. The "descent" is subtracting a small fraction of it from each weight.
Step size is the setting that matters
error
│╲ ╱
│ ╲ ●→●→●→● ╱ good learning rate: steady progress
│ ╲__________●____________╱
│
│╲ ●───────────────→● ╱ too large: it leaps past the bottom
│ ╲ ╱ and the error can blow up to NaN
│ ╲______________________╱That fraction is the learning rate, and it is the hyperparameter people spend the most time on. Too small and training crawls. Too large and each step overshoots the valley, the error climbs instead of falling, and you end up looking at a loss that has gone to NaN. A common starting point is 0.001 with the Adam optimizer, then adjust based on what the loss curve does in the first few hundred steps.
In practice nobody computes the slope over the entire dataset — that would be one step per full pass. Instead the slope is estimated from one batch at a time, which is called stochastic gradient descent. The estimate is noisy, and that noise turns out to help, because it shakes the model out of shallow dips that a perfectly smooth path would settle into. Optimizers like Adam add momentum and a per-weight step size on top of this basic idea.
Where to go next
- Full lesson: Gradient descent
- Related terms: backpropagation, loss-function, batch-size, hyperparameter