Adam (optimizer)
In one sentence Adam is the default optimizer for deep learning — momentum plus an automatic per-weight step size.
Updated
Adam is an optimizer that combines momentum with a separate, self-adjusting step size for every single weight.
Think of walking downhill in the dark with a torch. On ground you have crossed before and know is smooth, you stride. Where the torch shows rubble, you take small careful steps. One walking speed for all terrain — which is what plain SGD uses — is wrong almost everywhere.
Adam (short for adaptive moment estimation) keeps two running averages per weight. The first is the average gradient — the momentum part, giving direction with memory. The second is the average squared gradient, which measures how violently that weight's gradient swings. Weights with wild gradients get their steps shrunk; weights with steady, quiet gradients get relatively larger ones.
The practical result: Adam usually works decently with the default settings, on the first try. That robustness — not peak performance — is why it became the default. Well-tuned SGD with momentum can still match or beat it on vision tasks, but it demands more tuning of the learning-rate.
For transformers the standard variant is AdamW, which handles weight-decay correctly. The cost of all this bookkeeping is real: two extra numbers per parameter, which for a large model is many gigabytes of optimizer state.
Where to go next
- Full lesson: Optimization
- Related terms: sgd, momentum, learning-rate, weight-decay