AI glossary

Adam (optimizer)

In one sentence Adam is the default optimizer for deep learning — momentum plus an automatic per-weight step size.

By Updated

Adam is an optimizer that combines momentum with a separate, self-adjusting step size for every single weight.

Think of walking downhill in the dark with a torch. On ground you have crossed before and know is smooth, you stride. Where the torch shows rubble, you take small careful steps. One walking speed for all terrain — which is what plain SGD uses — is wrong almost everywhere.

Adam (short for adaptive moment estimation) keeps two running averages per weight. The first is the average gradient — the momentum part, giving direction with memory. The second is the average squared gradient, which measures how violently that weight's gradient swings. Weights with wild gradients get their steps shrunk; weights with steady, quiet gradients get relatively larger ones.

The practical result: Adam usually works decently with the default settings, on the first try. That robustness — not peak performance — is why it became the default. Well-tuned SGD with momentum can still match or beat it on vision tasks, but it demands more tuning of the learning-rate.

For transformers the standard variant is AdamW, which handles weight-decay correctly. The cost of all this bookkeeping is real: two extra numbers per parameter, which for a large model is many gigabytes of optimizer state.

Where to go next