AI glossary

Regularization

In one sentence Regularization is any extra pressure added during training that discourages a model from becoming complicated enough to memorise noise.

By Updated

Regularization is any technique that pushes a model toward staying simple, so it learns the pattern instead of memorising the training data.

Think of packing for a trip with a strict baggage limit. Without the limit you throw in everything "in case". With it, you are forced to pick what genuinely matters. Regularization is the baggage limit on a model: it makes complexity cost something, so the model spends its capacity only where the data truly demands it.

The most common form adds a penalty to the loss-function based on the size of the weights. L2 regularization (also called weight-decay) penalises the squares of the weights, shrinking them all gently toward zero. L1 regularization penalises absolute values, and tends to push many weights to exactly zero — a built-in feature selector.

Regularization is a family, not one trick. Dropout, early-stopping, smaller architectures, and data-augmentation all serve the same purpose: trading a little training accuracy for better behaviour on data the model has never seen.

The strength of the penalty is a hyperparameter. Too little and the model still overfits. Too much and it becomes too rigid to learn — the other side of the bias-variance-tradeoff.

Where to go next