AI glossary

Dropout

In one sentence Dropout randomly switches off a fraction of neurons during each training step, so no single neuron can be relied on too much.

By Updated

Dropout randomly turns off a fraction of a network's neurons at every training step, forcing the rest to cope without them.

Imagine a cricket team where, at every practice match, a random three players sit out. Nobody can build a strategy around one star batsman, because he might be missing today. The team is forced to develop many ways to win. Come the real match, everyone plays — and the team is more robust than one that always leaned on its star.

That is dropout exactly. During training, each neuron is zeroed out with some probability, often 10-50%. The network cannot let one neuron become the only detector of an important feature, so the knowledge spreads across many neurons. This fights overfitting, which is why dropout counts as regularization.

At inference time, dropout is switched off and all neurons participate. The outputs are scaled during training to compensate, so nothing needs adjusting later. In PyTorch this on/off switch is what model.train() and model.eval() control — forgetting model.eval() before evaluating is a classic bug that makes results noisy.

Dropout matters most in large fully-connected layers. Modern transformer training often uses little of it, because huge datasets and other regularizers do the same job.

Where to go next