Derivatives and gradients
A derivative says how much the result moves when you nudge one setting. A gradient collects that answer for every setting at once, and it is what training runs on.
- 18 min read
- 3 reading levels
- Published
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A derivative answers one question: if I nudge this one setting a little, how much does the result move?
Think about a shower with two taps, hot and cold. You turn the hot tap a small amount and the water gets warmer by some amount. You turn the cold tap the same small amount and it gets cooler by a different amount.
Each of those "how much did it change" answers is a derivative. You have been computing them with your hand for years.
One tap at a time
The taps matter by different amounts. On some showers the hot tap is fierce and a tiny turn scalds you. The cold tap might be feeble, needing half a turn to do anything.
Work out the effect of one tap while leaving the other alone. That is called a partial derivative: the change caused by moving one setting on its own.
turn the hot tap a little → water gets a LOT warmer
turn the cold tap a little → water gets a little coolerTwo settings, two answers, and they are different sizes. That difference is the useful information.
The gradient is the whole instruction sheet
Collect the answer for every tap into one list, in a fixed order. That list is the gradient.
The gradient tells you two things at once, and this is the part worth slowing down for.
Which way to turn each tap. A positive answer means turning that tap up pushes the result up.
How much each tap is worth. A large answer means that tap has a big effect right now. A near-zero answer means turning it does almost nothing at this setting.
Put together, the gradient points in the direction that changes the result fastest. In the previous lesson you stood on a foggy hillside and felt for the downhill direction. The gradient is that feeling, written down.
To go downhill, you take the gradient and go the opposite way. That is the whole update rule at the heart of training.
Why a model needs this
A model has settings, and a model has a score for how wrong it is. Training means turning the settings to make the wrongness smaller.
A small model has thousands of settings. A large one has billions. You cannot try them one by one, and nobody has that much time or electricity.
The gradient solves this. It tells you, for every single setting at once, which way to turn it and by roughly how much. One calculation, complete instructions.
how wrong we are → [ gradient ] → a small nudge for every setting
↓
turn them all, a little, at once
↓
slightly less wrongRepeat that a few hundred thousand times and you have a trained model.
The chain rule, with gears
Here is the idea that makes all of this possible on a big model.
Think about a bicycle with gears. You turn the pedal. The pedal turns the chain. The chain turns the rear wheel.
One turn of the pedal moves the chain by some amount. That chain movement turns the wheel by some further amount. The pedal's effect on the wheel is those two effects multiplied together.
pedal → chain → wheel
the pedal's effect on the wheel is
its effect on the chain, times the chain's effect on the wheelThat multiplying-along-the-chain rule is called the chain rule. A neural network is a very long chain: input, then layer, then layer, then layer, then a score.
To find how much the very first layer affects the final score, you multiply the effects along the whole chain. Doing that from the score backwards is called backpropagation, and there is a full lesson on it: backpropagation.
The honest part
Two real problems come straight out of that multiplying.
The message can fade. Play the whispering game down a line of people. By the end, the message is mush. If each link in the chain shrinks the effect a little, a long chain shrinks it to nothing. Early layers then get almost no instruction and stop learning. This is called the vanishing gradient problem, and it held deep learning back for years.
The message can explode. If each link amplifies instead, a long chain produces an absurd instruction. The model takes a wild step and the training run is ruined in one update.
Neither problem is fully solved. They are managed instead. Careful starting values help. So do particular layer designs. So does a cap on how large an instruction may be.
Where you have already used it
- Adjusting a music app's bass and treble until it sounds right, one slider at a time.
- Setting a ceiling fan's regulator, where the lower numbers barely differ and the top ones jump.
- Steering a car, where a small turn does much more at high speed than in a parking lot.
- Salting food, where the first pinch changes a lot and the fifth changes little.
Remember this
- A derivative says how much the result moves when you nudge one setting.
- The gradient collects that answer for every setting, giving direction and size at once.
- Effects multiply along a chain, which is how a change deep inside a model reaches the final score.
What to learn next
- Optimization — deciding how big a step to take.
- Backpropagation — the chain rule run backwards, in detail.
- Gradient descent — the loop that uses all of this.
Developer — Code and libraries.
Setup
pip install numpyThe main examples need NumPy alone. There is an optional PyTorch section at the end. It downloads a few hundred megabytes, so treat it as extra rather than required.
Calculus covered what a slope is and how finite differences behave. This page is about many settings at once, and about the chain rule.
Partial derivatives and the gradient
import numpy as np
def f(a, b):
return a * a + 3 * a * b + b * b * b
def grad_f(a, b):
return np.array([2 * a + 3 * b, # move a only
3 * a + 3 * b * b]) # move b only
point = np.array([2.0, 1.0])
g = grad_f(*point)
print("value at (2, 1) :", f(*point))
print("gradient :", g)
h = 1e-5
numeric = np.array([(f(point[0] + h, point[1]) - f(point[0] - h, point[1])) / (2 * h),
(f(point[0], point[1] + h) - f(point[0], point[1] - h)) / (2 * h)])
print("numeric check :", np.round(numeric, 6))
print("agree :", np.allclose(g, numeric, rtol=1e-6))
step = point - 0.05 * g # move against the gradient to go downhill
print("value after one downhill step:", round(f(*step), 4))value at (2, 1) : 11.0 gradient : [7. 9.] numeric check : [7. 9.] agree : True value after one downhill step: 5.6114
The gradient [7, 9] says b currently matters slightly more than a. Both entries are positive, so increasing either raises the value. To reduce it, subtract.
That subtraction took the value from 11.0 to 5.61 in one step. The step size 0.05 was a choice. A bad choice would have overshot, which is the subject of optimization.
The numeric check is a habit worth forming. Central differences on a hand-derived gradient catch sign errors and transposed indices in seconds. Use it whenever you write a gradient by hand.
Backpropagation, by hand, in one file
Here is a two-layer network on a single example, with every gradient derived and printed.
import numpy as np
x = np.array([1.0, 2.0]) # one example, two features
y = 5.0 # the answer we wanted
W1 = np.array([[0.5, -1.0], # 2 features in, 2 hidden neurons out
[1.5, 0.5]])
b1 = np.array([0.0, -0.5])
W2 = np.array([1.0, -2.0]) # 2 hidden in, 1 number out
b2 = 0.5
z1 = x @ W1 + b1
a1 = np.maximum(0.0, z1) # relu
z2 = a1 @ W2 + b2
loss = (z2 - y) ** 2
print("hidden before relu :", z1)
print("hidden after relu :", a1)
print("prediction :", z2, " target:", y)
print("loss :", loss)
dz2 = 2.0 * (z2 - y) # how the loss moves when the prediction moves
dW2 = a1 * dz2
db2 = dz2
da1 = W2 * dz2 # blame handed back to the hidden layer
dz1 = da1 * (z1 > 0) # relu passes blame only where it was awake
dW1 = np.outer(x, dz1)
db1 = dz1
print()
print("dL/dW2 :", dW2)
print("dL/db2 :", db2)
print("dL/dW1 :")
print(dW1)
print("dL/db1 :", db1)hidden before relu : [ 3.5 -0.5] hidden after relu : [3.5 0. ] prediction : 4.0 target: 5.0 loss : 1.0 dL/dW2 : [-7. -0.] dL/db2 : -2.0 dL/dW1 : [[-2. 0.] [-4. 0.]] dL/db1 : [-2. 0.]
Read the backward half as a walk down the same chain, in reverse. Each line converts blame from one stage into blame for the stage before it.
The zeros are the interesting part. The second hidden neuron had a pre-activation of -0.5, so ReLU output zero. Its whole column of gradients is zero: dL/dW1 column two, and the second entry of dL/db1.
That neuron receives no instruction from this example. It cannot learn from it.
Suppose a neuron stays negative for every example in your data. It then never updates again. That is the dead ReLU problem, and one reason Leaky ReLU exists. See activation functions.
np.outer(x, dz1) builds the weight gradient. Every weight W1[i, j] connects input i to hidden unit j, so its gradient is x[i] * dz1[j]. The outer product produces all four at once. That pattern — input on one side, incoming blame on the other — is every weight gradient in every network.
Verify it, do not trust it
import numpy as np
x = np.array([1.0, 2.0])
y = 5.0
W1 = np.array([[0.5, -1.0], [1.5, 0.5]])
b1 = np.array([0.0, -0.5])
W2 = np.array([1.0, -2.0])
b2 = 0.5
dW1 = np.array([[-2.0, 0.0], [-4.0, 0.0]]) # the hand-derived answer from above
h = 1e-5
def loss_with(W):
a = np.maximum(0.0, x @ W + b1)
return (a @ W2 + b2 - y) ** 2
for i in range(2):
for j in range(2):
up, dn = W1.copy(), W1.copy()
up[i, j] += h
dn[i, j] -= h
numeric = (loss_with(up) - loss_with(dn)) / (2 * h)
print(f"W1[{i},{j}] hand {dW1[i, j]:7.4f} numeric {numeric:10.6f}")W1[0,0] hand -2.0000 numeric -2.000000 W1[0,1] hand 0.0000 numeric 0.000000 W1[1,0] hand -4.0000 numeric -4.000000 W1[1,1] hand 0.0000 numeric 0.000000
Four matches. The hand-derived backward pass is correct, and now you know it rather than hope it.
Two rules for doing this on your own code. Run the check in float64, because float32 does not have the precision for the subtraction to survive. And compare with np.allclose(..., rtol=1e-5), never with ==.
The same thing, with autograd
pip install torchimport torch
x = torch.tensor([1.0, 2.0])
W1 = torch.tensor([[0.5, -1.0], [1.5, 0.5]], requires_grad=True)
b1 = torch.tensor([0.0, -0.5], requires_grad=True)
W2 = torch.tensor([1.0, -2.0], requires_grad=True)
b2 = torch.tensor(0.5, requires_grad=True)
loss = (torch.relu(x @ W1 + b1) @ W2 + b2 - 5.0) ** 2
loss.backward()
print("loss :", loss.item())
print("W1.grad :")
print(W1.grad)
print("W2.grad :", W2.grad)
print("b1.grad :", b1.grad)loss : 1.0
W1.grad :
tensor([[-2., 0.],
[-4., 0.]])
W2.grad : tensor([-7., -0.])
b1.grad : tensor([-2., 0.])Identical numbers, including the zeros from the dead neuron. loss.backward() did what you wrote out by hand. It would do the same for a network with a billion parameters.
requires_grad=True tells PyTorch to record every operation touching that tensor. The recording is a graph. backward() walks it in reverse, applying the chain rule at each node.
This is reverse-mode automatic differentiation. It is neither symbolic algebra nor finite differences. It is exact, and it costs about the same as one forward pass.
Common mistakes
1. Forgetting optimizer.zero_grad(). PyTorch accumulates gradients into .grad rather than replacing them. Skip the reset and step two silently uses the sum of two batches. Training gets worse for no visible reason.
2. Gradient checking in float32. The difference of two nearly equal float32 values has almost no correct digits left. Your correct gradient will look wrong.
3. Confusing the gradient with the loss. A large loss can come with a tiny gradient, on a flat plateau. A small loss can come with a large gradient. Log both; they answer different questions.
4. Reusing a stale forward pass. The backward pass needs the values from the forward pass at the current parameters. Change a weight, and you must run forward again before backward.
5. Reading a gradient of exactly zero as convergence. As the output above shows, zero often means a path was blocked, not that you have arrived.
Try it yourself
Change b1 in backprop.py to [0.0, 0.5]. That lifts the second hidden neuron above zero. Predict which entries of dL/dW1 stop being zero, then run it and check.
Then set W2 = np.array([1.0, 0.0]). Work out what happens to dL/dW1 and why, before you run it. A weight of zero blocks blame in exactly the way ReLU did.
What to learn next
- Optimization — turning these gradients into good step sizes.
- Backpropagation — the same algorithm across many layers and batches.
- PyTorch basics — autograd in a real training loop.
Researcher — Mathematics and papers.
Objects and shapes
For f: R**n -> R, the gradient collects the first partials:
grad_f(x) = [ df/dx_1, df/dx_2, ..., df/dx_n ].T shape (n,)For F: R**n -> R**m, the Jacobian collects all partials of all outputs:
J(x)[i][j] = dF_i / dx_j shape (m, n)For scalar f, the Hessian is the Jacobian of the gradient, H[i][j] = d**2 f / (dx_i dx_j), shape (n, n) and symmetric for C**2 functions.
The directional derivative along a unit vector u is D_u f(x) = grad_f(x).T @ u. By Cauchy-Schwarz it is maximised at u = grad_f / ||grad_f||. That is the formal statement that the gradient points in the direction of steepest ascent. The bound is tight only there.
One caveat that matters in practice: steepest descent is defined with respect to a norm. Gradient descent is steepest descent under the Euclidean norm. Natural gradient (Amari, 1998) is steepest descent under the Fisher information metric. Sign-based methods such as Adam approximate steepest descent under a max-norm-like geometry. "The" steepest direction is a choice of metric, not a fact.
The chain rule as matrix products
For a composition f = f_L ∘ ... ∘ f_1, the Jacobian is the product of layer Jacobians:
J_f(x) = J_{f_L}(h_{L-1}) @ ... @ J_{f_2}(h_1) @ J_{f_1}(x)with h_l the intermediate activations. Association is free mathematically and decisive computationally.
Evaluate right to left and you get forward mode: propagate J @ v for a direction v, one column at a time. Cost is O(n) passes for n inputs.
Evaluate left to right and you get reverse mode: propagate u.T @ J for a covector u, one row at a time. Cost is O(m) passes for m outputs.
Deep learning has one scalar loss and many parameters, so m = 1 and n in the billions. Reverse mode wins by a factor of n. This asymmetry is the single computational fact that makes deep learning feasible.
Baur-Strassen (1983) made the cost precise. The full gradient of a scalar function costs no more than a small constant factor times the function evaluation itself. That factor is commonly stated as about 3 to 4, and it does not grow with the number of inputs.
The primitive is the vector-Jacobian product (VJP). Frameworks never materialise a Jacobian; they register a VJP rule per operation. For Z = X @ W:
dL/dX = dL/dZ @ W.T
dL/dW = X.T @ dL/dZBoth are matrix multiplies of the same order as the forward pass, which is where the constant factor comes from.
Memory, the real constraint
Reverse mode must retain the forward activations needed by each VJP until the backward pass reaches them. Activation memory scales as O(B * sum_l d_l) for batch B, and it usually dominates parameter memory during training.
Gradient checkpointing (Chen et al., 2016) trades compute for memory. Store a subset of activations, and recompute the rest during the backward pass. With sqrt(L) checkpoints over L layers, memory drops to O(sqrt(L)). The price is roughly one extra forward pass. This is standard in large-model training.
Vanishing and exploding gradients
For a recurrent map with Jacobian J_t at each step, the gradient across T steps carries a product prod_{t} J_t. Bengio, Simard and Frasconi (1994) and Pascanu, Mikolov and Bengio (2013) analysed the consequence. If the spectral radius stays below 1, the product decays geometrically in T. Above 1, it grows geometrically. The stable case is a knife edge.
The mitigations in current use, and what each actually does:
- Initialisation. He (2015) sets
Var[W] = 2 / fan_infor ReLU, preserving activation variance across layers. Glorot and Bengio (2010) do the analogous thing for tanh. Both target the forward and backward variance simultaneously. - Residual connections (He et al., 2015) give the Jacobian an identity term, so the product has a path that neither decays nor explodes. This is why depth beyond about 20 layers became trainable at all.
- Normalisation layers rescale activations, which bounds the layer Jacobian in practice. Santurkar et al. (2018) argued the mechanism is smoothing of the loss landscape rather than the originally claimed internal covariate shift.
- Gradient clipping rescales when
||g|| > c:g <- g * c / ||g||. Zhang et al. (2020) showed clipping converges faster than fixed-step gradient descent under a relaxed smoothness condition that deep networks satisfy and standard L-smoothness does not. - Gating in LSTMs creates an additive cell path with near-identity Jacobian. See LSTM.
Higher order, and why it is rare
Newton's method uses d = -H_inverse @ grad, giving quadratic local convergence. It is unusable directly at scale. H is n x n, and neither forming nor inverting it is possible for n in the billions.
What is affordable is the Hessian-vector product, computable in about two gradient evaluations without forming H:
H @ v = d/d_eps [ grad_f(x + eps v) ] at eps = 0This is the Pearlmutter (1994) trick. It underpins Hessian-free optimisation, conjugate-gradient Newton methods, curvature-aware pruning and influence functions. K-FAC (Martens and Grosse, 2015) approximates the Fisher matrix in Kronecker-factored blocks, making a natural-gradient step tractable. In practice these rarely beat well-tuned Adam on wall-clock time for standard training, which is why first-order methods still dominate.
Practical numerics
Gradient checking should use central differences with h near 1e-5, in float64. Compare by relative error ||g_a - g_n|| / (||g_a|| + ||g_n||), with 1e-7 a reasonable pass threshold.
It fails legitimately near kinks. Nudging across a ReLU boundary changes which linear piece is active, so check at random points away from zero.
In mixed-precision training, gradients in float16 underflow to zero at magnitudes below roughly 6e-8. Loss scaling multiplies the loss by a large constant before the backward pass, and divides the gradients afterwards. bfloat16 largely removes the need, since it keeps the float32 exponent range while giving up significand bits.
References
- Rumelhart, D. E., Hinton, G. E., Williams, R. J. "Learning representations by back-propagating errors." Nature 323, 533–536, 1986.
- Baur, W., Strassen, V. "The complexity of partial derivatives." Theoretical Computer Science 22, 317–330, 1983.
- Griewank, A., Walther, A. Evaluating Derivatives, 2nd ed., SIAM, 2008. The standard reference on automatic differentiation.
- Baydin, A. G., Pearlmutter, B. A., Radul, A. A., Siskind, J. M. "Automatic Differentiation in Machine Learning: a Survey." JMLR 18, 2018. arXiv:1502.05767
- Pearlmutter, B. A. "Fast Exact Multiplication by the Hessian." Neural Computation 6(1), 147–160, 1994.
- Bengio, Y., Simard, P., Frasconi, P. "Learning long-term dependencies with gradient descent is difficult." IEEE Trans. Neural Networks 5(2), 157–166, 1994.
- Pascanu, R., Mikolov, T., Bengio, Y. "On the difficulty of training recurrent neural networks." ICML, 2013. arXiv:1211.5063
- Glorot, X., Bengio, Y. "Understanding the difficulty of training deep feedforward neural networks." AISTATS, 2010.
- He, K., Zhang, X., Ren, S., Sun, J. "Delving Deep into Rectifiers." ICCV, 2015. arXiv:1502.01852
- Chen, T., Xu, B., Zhang, C., Guestrin, C. "Training Deep Nets with Sublinear Memory Cost." arXiv:1604.06174, 2016.
- Zhang, J., He, T., Sra, S., Jadbabaie, A. "Why Gradient Clipping Accelerates Training." ICLR, 2020. arXiv:1905.11881
- Martens, J., Grosse, R. "Optimizing Neural Networks with Kronecker-factored Approximate Curvature." ICML, 2015. arXiv:1503.05671
- Amari, S. "Natural Gradient Works Efficiently in Learning." Neural Computation 10(2), 251–276, 1998.
What to learn next
- Optimization — convergence rates and adaptive step sizes.
- Backpropagation — the algorithm in full, layer by layer.
- LSTM — an architecture designed around the gradient product above.