Activation functions
An activation function is the bend placed after each layer, and without it every deep network collapses into a single straight-line model.
- 12 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
An activation function is a small bend added after each layer. Without it, a deep network can only ever draw straight lines.
Take a plastic ruler and draw a line along its edge. Now take a second ruler and continue from where you stopped, keeping it flat against the first. Add a third, a fourth, a tenth.
You still have one straight line. Stacking straight edges never produces a curve — you have to introduce a bend somewhere.
A neural network has exactly this problem. Each layer, on its own, does something ruler-like. Without a bend after each one, a hundred layers do no more work than one.
Why this had to be invented
Real questions are not straight lines.
"Is this mango ripe?" is not a matter of more softness always being better. A rock-hard mango is unripe, a slightly soft one is perfect, and a mushy one is rotten. The answer rises and then falls.
No straight line can rise and then fall. Something has to bend.
How it works
The bend sits between the layers, and it is applied to every number flowing through, one at a time.
inputs → [ mix them together ] → ( BEND ) → [ mix again ] → ( BEND ) → answer
layer 1 activation layer 2 activationRemove the bends and the two mixing steps quietly merge into one. This is not a matter of opinion. It is arithmetic you can check in three lines of code, and the Developer tab does that.
The bends people actually use
- ReLU — the most common one by a wide margin. Its rule is short: keep positive numbers as they are, and turn every negative number into zero. That single kink is enough.
- Sigmoid — squeezes any number into a value between zero and one. Useful at the very end, when you want the output to read as a confidence.
- Tanh — like sigmoid, but its output can be negative as well as positive.
- Softmax — used at the end when picking one label out of many. It turns a set of raw scores into percentages that add up to a hundred.
ReLU is the default for hidden layers. It is fast, and it does not fade out the correction signal the way sigmoid does deep in a network.
Where you have already seen this
Think of any photo filter that finds a face, or any assistant that hears its wake word. All of them have a ReLU, or a close relative, sitting between their layers. It is one of the most-executed pieces of code on earth.
What is honestly hard here
Choosing between ReLU, GELU and the others is not something you can reason out from first principles. Practitioners pick them mostly from experiment and convention.
That feels unsatisfying, and it is fair to feel that way. Much of deep learning is empirical — people tried things, some worked better, and the explanations came afterwards. Anyone who tells you there is a clean derivation for every choice is overselling it.
Remember this
- Without a bend after each layer, a deep network is no better than a single layer.
- ReLU is the standard bend for hidden layers: negatives become zero.
- Sigmoid and softmax belong at the end, where you want confidences.
What to learn next
- Backpropagation — where those derivatives in the table are actually used.
- Loss functions — the other half of the output layer decision.
- What is a neural network? — the structure these bends sit inside.
Developer — Code and libraries.
Setup
pip install numpyEvery common activation, plus proof the bend is required
import numpy as np
def relu(z):
return np.maximum(0.0, z)
def leaky_relu(z, slope=0.01):
return np.where(z > 0, z, slope * z)
def sigmoid(z):
return 1.0 / (1.0 + np.exp(-z))
def gelu(z):
# the tanh approximation, the one BERT and GPT-2 actually shipped with
return 0.5 * z * (1.0 + np.tanh(np.sqrt(2.0 / np.pi) * (z + 0.044715 * z ** 3)))
def row(name, values):
print(f"{name:9s}" + "".join(f"{v:8.3f}" for v in values))
x = np.array([-3.0, -1.0, -0.1, 0.0, 0.1, 1.0, 3.0])
row("input", x)
for name, fn in [("relu", relu), ("leaky", leaky_relu),
("sigmoid", sigmoid), ("tanh", np.tanh), ("gelu", gelu)]:
row(name, fn(x))
print("\n--- why a bend is needed ---")
A = np.array([[1.0, 2.0], [0.0, -1.0]]) # layer 1 weights
B = np.array([[2.0, 0.0], [1.0, 3.0]]) # layer 2 weights
v = np.array([[1.0, 3.0]])
row("2 linear", ((v @ A) @ B).ravel())
row("1 linear", (v @ (A @ B)).ravel())
row("relu mid", (relu(v @ A) @ B).ravel())
print("\n--- why sigmoid stalls learning ---")
for z in [0.0, 2.0, 6.0, 10.0]:
s = float(sigmoid(np.array(z)))
print(f"z={z:5.1f} sigmoid={s:.6f} slope={s * (1 - s):.8f}")
print("\n--- softmax ---")
def softmax(scores):
shifted = scores - scores.max() # subtracting the max stops exp() from overflowing
e = np.exp(shifted)
return e / e.sum()
small = np.array([2.0, 1.0, 3.0])
big = np.array([1002.0, 1001.0, 1003.0]) # same gaps, far larger numbers
row("small", softmax(small))
row("big", softmax(big))
print("each row sums to:", round(float(softmax(big).sum()), 6))input -3.000 -1.000 -0.100 0.000 0.100 1.000 3.000 relu 0.000 0.000 0.000 0.000 0.100 1.000 3.000 leaky -0.030 -0.010 -0.001 0.000 0.100 1.000 3.000 sigmoid 0.047 0.269 0.475 0.500 0.525 0.731 0.953 tanh -0.995 -0.762 -0.100 0.000 0.100 0.762 0.995 gelu -0.004 -0.159 -0.046 0.000 0.054 0.841 2.996 --- why a bend is needed --- 2 linear 1.000 -3.000 1 linear 1.000 -3.000 relu mid 2.000 0.000 --- why sigmoid stalls learning --- z= 0.0 sigmoid=0.500000 slope=0.25000000 z= 2.0 sigmoid=0.880797 slope=0.10499359 z= 6.0 sigmoid=0.997527 slope=0.00246651 z= 10.0 sigmoid=0.999955 slope=0.00004540 --- softmax --- small 0.245 0.090 0.665 big 0.245 0.090 0.665 each row sums to: 1.0
Reading the three experiments
The collapse proof. Rows 2 linear and 1 linear are identical, and that is the point. Two stacked weight matrices with nothing between them equal a single matrix A @ B. The relu mid row differs, because the bend broke the merge. This is the argument for activations, in three numbers.
The sigmoid stall. Watch the slope column. At z = 0 the slope is 0.25. At z = 10 it is 0.0000454 — about 5,500 times smaller. During training, corrections are multiplied by that slope on their way back through the network. A saturated sigmoid multiplies the correction by nearly zero, so the layers behind it stop learning. Stack several sigmoids and the signal reaching layer one is effectively gone. This is the vanishing gradient problem, and it is the reason ReLU replaced sigmoid in hidden layers.
Softmax stability. The small and big rows are identical. Softmax responds to the differences between scores, not their absolute size. Subtracting the maximum before exponentiating changes nothing mathematically, and it prevents np.exp(1003) from returning inf. Skip that line and you get RuntimeWarning: overflow encountered in exp followed by nan.
Common mistakes
Putting softmax on the output and then using a plain log loss. Compute cross-entropy from raw scores instead. In PyTorch, pass logits to nn.CrossEntropyLoss — it fuses the softmax internally for numerical stability. Applying softmax yourself and then CrossEntropyLoss applies it twice, and the model trains badly while raising no error at all.
Sigmoid in hidden layers. It works for two or three layers and then quietly stops working. Use ReLU or GELU in hidden layers, and keep sigmoid for a single binary output.
Dying ReLU. If a neuron's inputs push it negative for every example in the data, ReLU outputs zero, its slope is zero, and it receives no correction. It is dead for the rest of training. A large learning rate makes this much more likely. Fixes: lower the learning rate, switch to leaky_relu, or use GELU.
ReLU on the final layer of a regression model. It cannot output negative numbers. If your target can be negative — a temperature change, a profit-or-loss figure — use no activation at all on the last layer.
Try it yourself
Add swish to the table: z * sigmoid(z). Compare its column with gelu — they are close, which is a real result, not a coincidence. Then set slope=0.3 in leaky_relu and watch the negative side lift.
What to learn next
- Backpropagation — where those derivatives in the table are actually used.
- Loss functions — the other half of the output layer decision.
- What is a neural network? — the structure these bends sit inside.
Researcher — Mathematics and papers.
Definitions and derivatives
For scalar pre-activation $z$:
| Name | $\phi(z)$ | $\phi'(z)$ | Range |
|---|---|---|---|
| Sigmoid | $\sigma(z) = \dfrac{1}{1 + e^{-z}}$ | $\sigma(z)\left(1 - \sigma(z)\right)$ | $(0, 1)$ |
| Tanh | $\dfrac{e^{z} - e^{-z}}{e^{z} + e^{-z}}$ | $1 - \tanh^2(z)$ | $(-1, 1)$ |
| ReLU | $\max(0, z)$ | $\mathbb{1}[z > 0]$ | $[0, \infty)$ |
| Leaky ReLU | $\max(\alpha z, z)$ | $\mathbb{1}[z>0] + \alpha\,\mathbb{1}[z \le 0]$ | $(-\infty, \infty)$ |
| ELU | $z$ if $z>0$, else $\alpha(e^{z} - 1)$ | $1$ if $z>0$, else $\alpha e^{z}$ | $(-\alpha, \infty)$ |
| GELU | $z\,\Phi(z)$ | $\Phi(z) + z\,\varphi(z)$ | $(\approx -0.17, \infty)$ |
| SiLU / Swish | $z\,\sigma(z)$ | $\sigma(z)\left(1 + z(1 - \sigma(z))\right)$ | $(\approx -0.28, \infty)$ |
Where:
- $\alpha$ — a small positive constant, typically $0.01$ for leaky ReLU and $1.0$ for ELU.
- $\mathbb{1}[\cdot]$ — the indicator function, equal to $1$ when the condition holds and $0$ otherwise.
- $\Phi(z)$ — the CDF of the standard normal distribution.
- $\varphi(z)$ — the PDF of the standard normal distribution.
ReLU is not differentiable at $z = 0$. Frameworks assign a subgradient there — PyTorch returns $0$. This is harmless in practice, since exact zeros have measure zero in floating point.
Softmax and its Jacobian
For a score vector $z \in \mathbb{R}^{K}$:
$$ \text{softmax}(z)_i = \frac{e^{z_i}}{\sum_{j=1}^{K} e^{z_j}} $$
Its Jacobian is:
$$ \frac{\partial\, \text{softmax}(z)_i}{\partial z_j} = \text{softmax}(z)i \left( \delta{ij} - \text{softmax}(z)_j \right) $$
Where $\delta_{ij}$ is the Kronecker delta, equal to $1$ when $i = j$ and $0$ otherwise.
Softmax is shift-invariant: $\text{softmax}(z + c\mathbf{1}) = \text{softmax}(z)$ for any scalar $c$. Implementations exploit this by subtracting $\max_j z_j$, which bounds every exponent at $0$ and eliminates overflow. Use the log-sum-exp form when you need $\log \text{softmax}$:
$$ \log \text{softmax}(z)_i = z_i - \max_j z_j - \log \sum_{k} e^{z_k - \max_j z_j} $$
Saturation and gradient flow
The backward pass multiplies by $\phi'$ at every layer. For a depth-$L$ network the signal reaching layer $1$ carries a factor $\prod_{l=2}^{L} \phi'(z^{(l)})$.
Since $\max_z \sigma'(z) = 0.25$, a sigmoid network attenuates the gradient by at least $4^{-(L-1)}$ even in the best case. At $L = 10$ that is a factor of $4 \times 10^{-6}$. This is the vanishing gradient problem, analysed by Hochreiter (1991) and Bengio, Simard and Frasconi (1994).
ReLU has $\phi' \in {0, 1}$, so it neither shrinks nor amplifies the gradient on its active path. The cost is the dying ReLU failure mode: once a unit's pre-activation is negative across the whole data distribution, its gradient is identically zero and it never recovers. Lu et al. (2019), Dying ReLU and initialization, show the probability of a network being born dead grows with depth for certain initialisations.
Initialisation is coupled to the activation
Variance-preserving initialisation depends on which activation follows it. For a layer with $n_{\text{in}}$ inputs and $n_{\text{out}}$ outputs:
- Glorot / Xavier (2010), for symmetric saturating activations such as tanh: $\operatorname{Var}(W) = \dfrac{2}{n_{\text{in}} + n_{\text{out}}}$.
- He (2015), for ReLU: $\operatorname{Var}(W) = \dfrac{2}{n_{\text{in}}}$.
The factor of $2$ in He initialisation compensates for ReLU zeroing roughly half its inputs, which halves the variance of the output. Using Glorot with ReLU in a deep network causes activations to shrink layer by layer, and training stalls. This pairing is a genuine source of silent bugs.
State of the art
- ReLU — Nair and Hinton (2010), Rectified linear units improve restricted Boltzmann machines, with Krizhevsky et al. (2012) demonstrating it at scale in AlexNet.
- GELU — Hendrycks and Gimpel (2016), Gaussian error linear units. Standard in BERT, GPT-2 and most transformer encoders.
- Swish / SiLU — Ramachandran et al. (2017), Searching for activation functions, found by automated search. Note that SiLU was described earlier by Elfwing et al. (2017).
- GLU variants — Shazeer (2020), GLU variants improve transformer. SwiGLU is now the default feedforward activation in Llama, PaLM and most current large language models. It uses a gated form $\text{SwiGLU}(x) = \text{Swish}(xW) \odot (xV)$, where $\odot$ is elementwise multiplication, and needs three weight matrices instead of two — so implementations shrink the hidden dimension to $\tfrac{2}{3}$ of the usual size to keep the parameter count matched.
The empirical gaps between modern choices are small — often a fraction of a percent. Architecture, data quality and scale dominate. Treat activation choice as a low-priority hyperparameter unless you are training at frontier scale.
What to learn next
- Backpropagation — where those derivatives in the table are actually used.
- Loss functions — the other half of the output layer decision.
- What is a neural network? — the structure these bends sit inside.