AI glossary

Activation function

In one sentence An activation function is the small bending step applied to each neuron's output, and it is what lets a deep network learn curved patterns instead of straight lines.

By Updated

An activation function is a small non-linear step applied to each neuron's output, so the network can learn curved, complicated patterns.

Think of a water tap on a pipe. The pipe carries whatever comes in; the tap decides how much actually gets through. Some taps shut off completely below a certain pressure. Some open gradually. Each neuron in a network has one of these taps sitting on its output, and the shape of the tap changes what the whole network is able to learn.

Without that bend, depth is worthless. Stacking straight lines gives you another straight line, so a hundred layers of pure multiplication and addition would have exactly the power of one layer. The activation function is the one ingredient that makes "deep" mean something.

ReLU, the one you will meet first

ReLU keeps positive numbers and flattens everything negative to zero.

input:   -3.0   -0.5    0.0    1.2    4.7
ReLU:     0.0    0.0    0.0    1.2    4.7

That is the whole rule, and it is the default in most vision and many language models because it is cheap and it trains well. Its known weakness is the "dying ReLU". If a neuron's input is negative for every example in the data, its output is always zero, its slope is zero, and it stops learning for good. That is why variants like LeakyReLU and GELU exist. GELU is the common choice inside transformers.

A different family lives at the end of a network: sigmoid squeezes one number into 0-to-1 for yes/no answers, and softmax turns a row of scores into probabilities that add to one. Those are output activations, chosen to match the loss function, and in PyTorch you usually leave them out of the model and let CrossEntropyLoss apply them internally.

Where to go next