AI glossary

Batch size

In one sentence Batch size is how many training examples the model looks at before it updates its weights once.

By Updated

Batch size is how many training examples the model sees before it updates its weights one time.

Rolling chapatis is the right picture. You could roll one, put it on the tawa, wait, and roll the next — safe, but slow. You could roll fifty first, then cook them all — fast, but you need a huge counter to hold fifty. Batch size is how many you roll before moving to the tawa. The counter is your GPU memory.

The model does not update after every single example. It works through a batch, averages the error across all of them, and takes one step. A batch of 32 means 32 predictions, one average error, one adjustment to the weights.

What the number changes

50,000 training images

batch size 100  →  500 weight updates per epoch, small memory, noisy steps
batch size 500  →  100 weight updates per epoch, large memory, smooth steps

A larger batch gives a steadier estimate of the right direction, uses the GPU more efficiently, and needs more memory. A smaller batch is noisier, and that noise is not purely bad — it acts a little like a nudge out of bad spots, and small batches often generalise well. What matters more than the exact value is that batch size and learning rate move together. Double the batch and the same learning rate takes effectively smaller steps per example. So people commonly scale the learning rate up when they scale the batch up.

Batch size is also the first number to lower when training dies with CUDA out of memory. If a smaller batch hurts your results, gradient accumulation lets you keep the effective batch size while only holding a few examples in memory at a time.

Where to go next