AI glossary

Layer normalization

In one sentence Layer normalization rescales each example's features using that example alone, which is why transformers prefer it to batch norm.

By Updated

Layer normalization rescales the features of each example using only that example's own numbers, with no dependence on the rest of the batch.

Grading exams makes the difference concrete. Batch-normalization grades on a curve: your adjusted score depends on how everyone else in the room did. Layer norm grades each paper against itself: each answer is scaled relative to that student's own average. No classmates needed — the result is identical whether one paper is graded or a thousand.

Mechanically, for one example it takes all the feature values at a layer, subtracts their mean, divides by their spread, then applies a learned scale and shift. Same recipe as batch norm; the difference is purely which direction the averaging runs.

That independence is why transformers use it. Language batches contain sequences of wildly different lengths, batch sizes vary between training and serving, and generation often runs one token at a time — situations where batch statistics are meaningless or unavailable. Layer norm behaves identically in training and inference, with no running averages to maintain.

In a transformer you will find it wrapped around every attention and feed-forward block, usually in the "pre-norm" position that makes very deep stacks trainable. A popular simplification called RMSNorm drops the mean-subtraction step and is used by many recent LLMs, including the Llama family.

Where to go next