Momentum
In one sentence Momentum makes each training step a running blend of past steps, so updates build speed in consistent directions and cancel out zigzag.
Updated
Momentum makes each weight update a running average of recent updates, so training accelerates in directions that stay consistent and damps directions that flip.
A ball rolling down a hilly slope does not stop at every pebble. It carries speed from the slope it has already descended, coasting over small bumps and settling only in a genuinely deep valley. A walker with no memory — plain SGD — reacts to every pebble underfoot.
Mechanically, the optimizer keeps a velocity for every weight: a decayed sum of past gradients, typically keeping about 90% of the previous velocity each step. If gradients keep pointing the same way, velocity builds and progress speeds up. If gradients alternate signs — the zigzag you get in narrow, steep valleys — they cancel inside the average, and the wobble smooths out.
without momentum: \/\/\/\/\/\/\/ bounces across the valley walls
with momentum : ————————→ averages out the bounce, moves along the valleyThe cost is one extra number stored per parameter and one more hyperparameter to set. Nearly every modern optimizer keeps the idea: Adam is, at heart, momentum plus a per-weight step size.
Where to go next
- Full lesson: Optimization
- Related terms: sgd, adam, learning-rate, gradient-descent