SGD (stochastic gradient descent)
In one sentence SGD updates the model after each small random batch of data instead of after the whole dataset, trading noisy steps for speed.
Updated
Stochastic gradient descent, or SGD, is gradient-descent that updates the weights after each small random batch of examples, rather than after seeing the entire dataset.
Suppose you are cooking dal for two hundred guests and want the salt right. The "full" method tastes every single plate before touching the salt — accurate, and absurdly slow. The stochastic method tastes one spoonful, adjusts, stirs, tastes another. Each spoonful is an imperfect estimate, but you adjust hundreds of times in the time one full tasting would take. Stochastic is the technical word for "based on a random sample".
The noise in each step is not purely a cost. Full-dataset steps follow one smooth path and can settle into the nearest dip in the loss surface. Noisy steps jitter, and that jitter helps the model escape shallow dips and often find flatter, better solutions.
full-batch descent : one careful step per pass over all data
SGD : hundreds of rough steps in the same timeIn practice, "SGD" nearly always means mini-batch SGD: batches of 32 or 256 examples, not single ones, because GPUs process a batch in parallel almost for free. The batch-size sets where you sit between noisy and slow. Plain SGD is often paired with momentum, and Adam builds further on the same idea.
Where to go next
- Full lesson: Gradient descent
- Related terms: gradient-descent, momentum, adam, batch-size