AI glossary

Batch normalization

In one sentence Batch normalization rescales each layer's outputs using the current batch's statistics, keeping the numbers in a healthy range so training stays stable.

By Updated

Batch normalization rescales the outputs of a layer so that, across the current batch, they have a steady mean and spread — keeping every layer's numbers in a range the next layer can handle.

Think of a relay of loudspeakers passing a message across a field. If one speaker shouts, the next receives distortion; if one whispers, the next receives nothing. The fix is a volume knob at every station, retuned constantly, so each speaker receives sound at the same comfortable level. Batch norm is that knob between layers.

Concretely, for each feature it takes the values across the current batch-size worth of examples, subtracts their mean, and divides by their spread (the standard deviation, a measure of how scattered values are). Then two learned parameters per feature let the network scale and shift the result — so normalisation helps stability without limiting what the layer can express.

The benefits are practical: training tolerates higher learning-rates, converges faster, and depends less on lucky weight initialisation. It also adds a mild regularising effect, since each example's normalisation depends on which batch-mates it was grouped with.

The dependence on the batch is also its weakness. At inference time there is no batch, so running averages saved during training are used instead — another behaviour switched by model.eval(). And with very small batches the statistics get noisy, which is one reason transformers use layer-normalization instead.

Where to go next