Deep Learning

Activation functions

An activation function is the bend placed after each layer, and without it every deep network collapses into a single straight-line model.

Read these first

On this page 7
  1. Why this had to be invented
  2. How it works
  3. The bends people actually use
  4. Where you have already seen this
  5. What is honestly hard here
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

An activation function is a small bend added after each layer. Without it, a deep network can only ever draw straight lines.

Take a plastic ruler and draw a line along its edge. Now take a second ruler and continue from where you stopped, keeping it flat against the first. Add a third, a fourth, a tenth.

You still have one straight line. Stacking straight edges never produces a curve — you have to introduce a bend somewhere.

A neural network has exactly this problem. Each layer, on its own, does something ruler-like. Without a bend after each one, a hundred layers do no more work than one.

Why this had to be invented

Real questions are not straight lines.

"Is this mango ripe?" is not a matter of more softness always being better. A rock-hard mango is unripe, a slightly soft one is perfect, and a mushy one is rotten. The answer rises and then falls.

No straight line can rise and then fall. Something has to bend.

How it works

The bend sits between the layers, and it is applied to every number flowing through, one at a time.

  inputs → [ mix them together ] → ( BEND ) → [ mix again ] → ( BEND ) → answer
              layer 1              activation    layer 2      activation

Remove the bends and the two mixing steps quietly merge into one. This is not a matter of opinion. It is arithmetic you can check in three lines of code, and the Developer tab does that.

The bends people actually use

  • ReLU — the most common one by a wide margin. Its rule is short: keep positive numbers as they are, and turn every negative number into zero. That single kink is enough.
  • Sigmoid — squeezes any number into a value between zero and one. Useful at the very end, when you want the output to read as a confidence.
  • Tanh — like sigmoid, but its output can be negative as well as positive.
  • Softmax — used at the end when picking one label out of many. It turns a set of raw scores into percentages that add up to a hundred.

ReLU is the default for hidden layers. It is fast, and it does not fade out the correction signal the way sigmoid does deep in a network.

Where you have already seen this

Think of any photo filter that finds a face, or any assistant that hears its wake word. All of them have a ReLU, or a close relative, sitting between their layers. It is one of the most-executed pieces of code on earth.

What is honestly hard here

Choosing between ReLU, GELU and the others is not something you can reason out from first principles. Practitioners pick them mostly from experiment and convention.

That feels unsatisfying, and it is fair to feel that way. Much of deep learning is empirical — people tried things, some worked better, and the explanations came afterwards. Anyone who tells you there is a clean derivation for every choice is overselling it.

Remember this

  • Without a bend after each layer, a deep network is no better than a single layer.
  • ReLU is the standard bend for hidden layers: negatives become zero.
  • Sigmoid and softmax belong at the end, where you want confidences.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Every common activation, plus proof the bend is required

activations.py
import numpy as np


def relu(z):
    return np.maximum(0.0, z)


def leaky_relu(z, slope=0.01):
    return np.where(z > 0, z, slope * z)


def sigmoid(z):
    return 1.0 / (1.0 + np.exp(-z))


def gelu(z):
    # the tanh approximation, the one BERT and GPT-2 actually shipped with
    return 0.5 * z * (1.0 + np.tanh(np.sqrt(2.0 / np.pi) * (z + 0.044715 * z ** 3)))


def row(name, values):
    print(f"{name:9s}" + "".join(f"{v:8.3f}" for v in values))


x = np.array([-3.0, -1.0, -0.1, 0.0, 0.1, 1.0, 3.0])
row("input", x)
for name, fn in [("relu", relu), ("leaky", leaky_relu),
                 ("sigmoid", sigmoid), ("tanh", np.tanh), ("gelu", gelu)]:
    row(name, fn(x))

print("\n--- why a bend is needed ---")
A = np.array([[1.0, 2.0], [0.0, -1.0]])     # layer 1 weights
B = np.array([[2.0, 0.0], [1.0, 3.0]])      # layer 2 weights
v = np.array([[1.0, 3.0]])

row("2 linear", ((v @ A) @ B).ravel())
row("1 linear", (v @ (A @ B)).ravel())
row("relu mid", (relu(v @ A) @ B).ravel())

print("\n--- why sigmoid stalls learning ---")
for z in [0.0, 2.0, 6.0, 10.0]:
    s = float(sigmoid(np.array(z)))
    print(f"z={z:5.1f}   sigmoid={s:.6f}   slope={s * (1 - s):.8f}")

print("\n--- softmax ---")


def softmax(scores):
    shifted = scores - scores.max()         # subtracting the max stops exp() from overflowing
    e = np.exp(shifted)
    return e / e.sum()


small = np.array([2.0, 1.0, 3.0])
big = np.array([1002.0, 1001.0, 1003.0])    # same gaps, far larger numbers
row("small", softmax(small))
row("big", softmax(big))
print("each row sums to:", round(float(softmax(big).sum()), 6))
Output
input      -3.000  -1.000  -0.100   0.000   0.100   1.000   3.000
relu        0.000   0.000   0.000   0.000   0.100   1.000   3.000
leaky      -0.030  -0.010  -0.001   0.000   0.100   1.000   3.000
sigmoid     0.047   0.269   0.475   0.500   0.525   0.731   0.953
tanh       -0.995  -0.762  -0.100   0.000   0.100   0.762   0.995
gelu       -0.004  -0.159  -0.046   0.000   0.054   0.841   2.996

--- why a bend is needed ---
2 linear    1.000  -3.000
1 linear    1.000  -3.000
relu mid    2.000   0.000

--- why sigmoid stalls learning ---
z=  0.0   sigmoid=0.500000   slope=0.25000000
z=  2.0   sigmoid=0.880797   slope=0.10499359
z=  6.0   sigmoid=0.997527   slope=0.00246651
z= 10.0   sigmoid=0.999955   slope=0.00004540

--- softmax ---
small       0.245   0.090   0.665
big         0.245   0.090   0.665
each row sums to: 1.0

Reading the three experiments

The collapse proof. Rows 2 linear and 1 linear are identical, and that is the point. Two stacked weight matrices with nothing between them equal a single matrix A @ B. The relu mid row differs, because the bend broke the merge. This is the argument for activations, in three numbers.

The sigmoid stall. Watch the slope column. At z = 0 the slope is 0.25. At z = 10 it is 0.0000454 — about 5,500 times smaller. During training, corrections are multiplied by that slope on their way back through the network. A saturated sigmoid multiplies the correction by nearly zero, so the layers behind it stop learning. Stack several sigmoids and the signal reaching layer one is effectively gone. This is the vanishing gradient problem, and it is the reason ReLU replaced sigmoid in hidden layers.

Softmax stability. The small and big rows are identical. Softmax responds to the differences between scores, not their absolute size. Subtracting the maximum before exponentiating changes nothing mathematically, and it prevents np.exp(1003) from returning inf. Skip that line and you get RuntimeWarning: overflow encountered in exp followed by nan.

Common mistakes

Putting softmax on the output and then using a plain log loss. Compute cross-entropy from raw scores instead. In PyTorch, pass logits to nn.CrossEntropyLoss — it fuses the softmax internally for numerical stability. Applying softmax yourself and then CrossEntropyLoss applies it twice, and the model trains badly while raising no error at all.

Sigmoid in hidden layers. It works for two or three layers and then quietly stops working. Use ReLU or GELU in hidden layers, and keep sigmoid for a single binary output.

Dying ReLU. If a neuron's inputs push it negative for every example in the data, ReLU outputs zero, its slope is zero, and it receives no correction. It is dead for the rest of training. A large learning rate makes this much more likely. Fixes: lower the learning rate, switch to leaky_relu, or use GELU.

ReLU on the final layer of a regression model. It cannot output negative numbers. If your target can be negative — a temperature change, a profit-or-loss figure — use no activation at all on the last layer.

Try it yourself

Add swish to the table: z * sigmoid(z). Compare its column with gelu — they are close, which is a real result, not a coincidence. Then set slope=0.3 in leaky_relu and watch the negative side lift.

What to learn next

Researcher — Mathematics and papers.

Definitions and derivatives

For scalar pre-activation $z$:

Name$\phi(z)$$\phi'(z)$Range
Sigmoid$\sigma(z) = \dfrac{1}{1 + e^{-z}}$$\sigma(z)\left(1 - \sigma(z)\right)$$(0, 1)$
Tanh$\dfrac{e^{z} - e^{-z}}{e^{z} + e^{-z}}$$1 - \tanh^2(z)$$(-1, 1)$
ReLU$\max(0, z)$$\mathbb{1}[z > 0]$$[0, \infty)$
Leaky ReLU$\max(\alpha z, z)$$\mathbb{1}[z>0] + \alpha\,\mathbb{1}[z \le 0]$$(-\infty, \infty)$
ELU$z$ if $z>0$, else $\alpha(e^{z} - 1)$$1$ if $z>0$, else $\alpha e^{z}$$(-\alpha, \infty)$
GELU$z\,\Phi(z)$$\Phi(z) + z\,\varphi(z)$$(\approx -0.17, \infty)$
SiLU / Swish$z\,\sigma(z)$$\sigma(z)\left(1 + z(1 - \sigma(z))\right)$$(\approx -0.28, \infty)$

Where:

  • $\alpha$ — a small positive constant, typically $0.01$ for leaky ReLU and $1.0$ for ELU.
  • $\mathbb{1}[\cdot]$ — the indicator function, equal to $1$ when the condition holds and $0$ otherwise.
  • $\Phi(z)$ — the CDF of the standard normal distribution.
  • $\varphi(z)$ — the PDF of the standard normal distribution.

ReLU is not differentiable at $z = 0$. Frameworks assign a subgradient there — PyTorch returns $0$. This is harmless in practice, since exact zeros have measure zero in floating point.

Softmax and its Jacobian

For a score vector $z \in \mathbb{R}^{K}$:

$$ \text{softmax}(z)_i = \frac{e^{z_i}}{\sum_{j=1}^{K} e^{z_j}} $$

Its Jacobian is:

$$ \frac{\partial\, \text{softmax}(z)_i}{\partial z_j} = \text{softmax}(z)i \left( \delta{ij} - \text{softmax}(z)_j \right) $$

Where $\delta_{ij}$ is the Kronecker delta, equal to $1$ when $i = j$ and $0$ otherwise.

Softmax is shift-invariant: $\text{softmax}(z + c\mathbf{1}) = \text{softmax}(z)$ for any scalar $c$. Implementations exploit this by subtracting $\max_j z_j$, which bounds every exponent at $0$ and eliminates overflow. Use the log-sum-exp form when you need $\log \text{softmax}$:

$$ \log \text{softmax}(z)_i = z_i - \max_j z_j - \log \sum_{k} e^{z_k - \max_j z_j} $$

Saturation and gradient flow

The backward pass multiplies by $\phi'$ at every layer. For a depth-$L$ network the signal reaching layer $1$ carries a factor $\prod_{l=2}^{L} \phi'(z^{(l)})$.

Since $\max_z \sigma'(z) = 0.25$, a sigmoid network attenuates the gradient by at least $4^{-(L-1)}$ even in the best case. At $L = 10$ that is a factor of $4 \times 10^{-6}$. This is the vanishing gradient problem, analysed by Hochreiter (1991) and Bengio, Simard and Frasconi (1994).

ReLU has $\phi' \in {0, 1}$, so it neither shrinks nor amplifies the gradient on its active path. The cost is the dying ReLU failure mode: once a unit's pre-activation is negative across the whole data distribution, its gradient is identically zero and it never recovers. Lu et al. (2019), Dying ReLU and initialization, show the probability of a network being born dead grows with depth for certain initialisations.

Initialisation is coupled to the activation

Variance-preserving initialisation depends on which activation follows it. For a layer with $n_{\text{in}}$ inputs and $n_{\text{out}}$ outputs:

  • Glorot / Xavier (2010), for symmetric saturating activations such as tanh: $\operatorname{Var}(W) = \dfrac{2}{n_{\text{in}} + n_{\text{out}}}$.
  • He (2015), for ReLU: $\operatorname{Var}(W) = \dfrac{2}{n_{\text{in}}}$.

The factor of $2$ in He initialisation compensates for ReLU zeroing roughly half its inputs, which halves the variance of the output. Using Glorot with ReLU in a deep network causes activations to shrink layer by layer, and training stalls. This pairing is a genuine source of silent bugs.

State of the art

  • ReLU — Nair and Hinton (2010), Rectified linear units improve restricted Boltzmann machines, with Krizhevsky et al. (2012) demonstrating it at scale in AlexNet.
  • GELU — Hendrycks and Gimpel (2016), Gaussian error linear units. Standard in BERT, GPT-2 and most transformer encoders.
  • Swish / SiLU — Ramachandran et al. (2017), Searching for activation functions, found by automated search. Note that SiLU was described earlier by Elfwing et al. (2017).
  • GLU variants — Shazeer (2020), GLU variants improve transformer. SwiGLU is now the default feedforward activation in Llama, PaLM and most current large language models. It uses a gated form $\text{SwiGLU}(x) = \text{Swish}(xW) \odot (xV)$, where $\odot$ is elementwise multiplication, and needs three weight matrices instead of two — so implementations shrink the hidden dimension to $\tfrac{2}{3}$ of the usual size to keep the parameter count matched.

The empirical gaps between modern choices are small — often a fraction of a percent. Architecture, data quality and scale dominate. Treat activation choice as a low-priority hyperparameter unless you are training at frontier scale.

What to learn next