AI glossary

Optimizer

In one sentence The optimizer is the rule that decides how to change each weight after the gradients are computed.

By Updated

The optimizer is the component that takes the gradients and decides exactly how much to change every weight.

Training has a clean division of labour. Backpropagation is the surveyor: it measures which direction is downhill for every weight. The optimizer is the driver: it decides how far to move, whether to carry speed from previous steps, and whether some wheels should turn more than others. Same map, different driving styles, very different journeys.

The family tree is short and worth knowing:

gradient descent    take a plain step downhill
+ mini-batches   →  SGD: step after each small random batch
+ velocity       →  SGD with momentum: smooth the zigzag, coast through bumps
+ per-weight step→  Adam / AdamW: each weight gets its own adaptive step size

All of them are steered by the learning-rate, and each keeps different amounts of state. Plain SGD stores nothing extra; Adam stores two numbers per parameter, which is why optimizer state dominates memory when training large models.

Choosing one is less dramatic than beginners fear. AdamW is the safe default for transformers and most new projects. SGD with momentum remains competitive for vision models when you can afford to tune it.

Where to go next