Mathematics for AI

Calculus

Calculus is the maths of change — how fast something is changing right now, and how much piles up over time. Training a model is the first half, repeated.

On this page 9
  1. Why you should care
  2. The awkward idea at the centre
  3. Zoom in far enough and every curve is straight
  4. What this has to do with AI
  5. The other dial
  6. The honest part
  7. Where you have already seen it
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Calculus is the maths of change: how fast something is changing right now, and how much piles up over time.

Think about sitting in a bus and looking at the dashboard. Two dials. One says how fast you are moving at this moment. The other counts up the kilometres you have covered.

Those two dials are the two halves of calculus, and you have watched both of them your whole life. The rest is care about what "at this moment" means.

Why you should care

Training an AI model is one move, repeated. Work out which small change makes the model less wrong. Then make it.

"Which small change" is a question about rate of change. It is the speedometer half, and it runs millions of times inside every training job.

Why does training sometimes blow up? Why does a loss curve flatten and stop? Why do some models refuse to learn at all? This is the lesson underneath all three.

The awkward idea at the centre

Speed is distance divided by time. So what is your speed at a single instant?

At an instant, no time passes and no distance is covered. Nothing divided by nothing is not an answer. This bothered serious people for a very long time.

The escape is to stop asking about an instant, and ask about shorter and shorter stretches instead.

Measure your speed over an hour. Then over a minute. Then over a second. Then over a blink. The answers settle down toward one number, and that number is what we mean by speed right now.

This settling-down is called a limit — the value something creeps toward as you keep shrinking the gap. It is the one genuinely new idea in the whole subject.

Zoom in far enough and every curve is straight

Here is the picture that makes it click.

   the whole road, from far away        the same road, zoomed right in

        ____                                       ______
       /    \___                                  /
      /         \                          ______/
     /           \___

A winding road looks curved from a distance. Stand on it and the patch under your feet is flat and tilted one way.

That tilt is the slope — how steeply things rise or fall right where you stand. Calculus is the machinery for finding that tilt anywhere on the curve, without walking the whole road.

What this has to do with AI

A model is graded by how wrong it is. Picture that wrongness as the height of a landscape.

   very wrong   \                              /
                 \                            /
                  \___                    ___/
                      \____          ____/
                           \___  ___/
   least wrong                \__/
                            the setting you want

You are standing somewhere on that landscape, in thick fog. You cannot see the bottom. What you can feel is the tilt of the ground under your feet.

The tilt tells you which way is downhill. Take a step that way. Feel again. Step again.

That is training. Every model on this website was trained by that loop, and the tilt at each step comes from calculus. The next lesson, derivatives and gradients, is entirely about how the tilt is worked out.

The other dial

The odometer half is called integration: adding up a huge number of tiny pieces to get a total.

Picture pouring water into a tank. The tap controls the rate. The level in the tank is everything that has flowed in so far. Turn the tap up and the level climbs faster.

The surprising discovery, and the one that made calculus famous, is that the two dials are opposites. The tap rate tells you how the level changes. The level is the tap rate accumulated.

In AI, this half shows up quietly rather than loudly. Chances of things happening are areas under curves. Averages over a whole population are the same kind of sum. You will use it far less often than the first half.

The honest part

Two things are worth saying plainly.

You will not do calculus by hand. Libraries compute rates of change automatically, for models with billions of settings. Nobody differentiates those by hand, and nobody expects you to.

Not every curve is smooth. Some have sharp corners, like the point of a roof. At the corner itself, "the tilt" has no single answer — it depends on which side you approach from.

That matters because one of the most widely used pieces in deep learning has exactly such a corner. It works anyway, and the reason it works is a real and slightly untidy story. The Developer block shows the corner in code.

Where you have already seen it

  • Your phone's "two hours of battery left" — a rate of change, turned into a prediction.
  • A speedometer, which is a rate of change with a needle attached.
  • The fuel gauge, dropping fast in traffic and slowly on the highway.
  • Interest on a loan, piling up a little at a time.

Remember this

  • Calculus has two halves: how fast something changes now, and how much accumulates over time.
  • The trick behind both is shrinking a gap until the answer settles: a limit.
  • Training a model means feeling the tilt of a landscape of wrongness and stepping downhill.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Three short programs. The first measures a slope numerically and shows exactly where that method breaks. The second accumulates an area. The third finds the corner in ReLU.

Measuring a slope, and watching it fail

The derivative is the slope of a function at a point. The definition shrinks a step h toward zero. On a computer, h cannot go all the way there.

slope.py
def f(x):
    return x * x            # the exact slope at x is 2x, so 6.0 at x = 3

x0 = 3.0
print(f"{'step h':>10} {'forward':>18} {'central':>18}")
for h in (1e-1, 1e-3, 1e-6, 1e-10, 1e-14):
    forward = (f(x0 + h) - f(x0)) / h
    central = (f(x0 + h) - f(x0 - h)) / (2 * h)
    print(f"{h:>10.0e} {forward:>18.10f} {central:>18.10f}")
Output
    step h            forward            central
     1e-01       6.1000000000       6.0000000000
     1e-03       6.0010000000       6.0000000000
     1e-06       6.0000010009       6.0000000008
     1e-10       6.0000004964       6.0000004964
     1e-14       6.2172489379       6.1284310959

The true answer is 6.0. Read the table from the top and two separate effects fight each other.

Shrinking h helps, then hurts. The forward estimate improves down to 1e-6. After that it gets worse. By 1e-14 it is wrong in the first decimal place. Every step of that is expected, and none of it is a bug.

The cause is catastrophic cancellation. When h is tiny, f(x0 + h) and f(x0) agree in almost every digit. Subtracting them destroys those digits and leaves mostly rounding error. Dividing by a tiny h then magnifies what is left.

The practical sweet spot for the forward difference is around the square root of machine epsilon, about 1e-8 in float64. Below that you are amplifying noise.

The central difference is far better. At h = 1e-1 it is already exact for this function, because the errors on either side cancel. Its error shrinks with h**2 rather than h. When you need a numerical derivative, use the central form.

There is a bigger point here. Numerical differentiation costs one function evaluation per parameter. A model with a million parameters would need a million forward passes for one gradient step. This is why nobody trains with finite differences, and why automatic differentiation exists.

Finite differences are still worth knowing, for one job: checking that a hand-written gradient is correct. That check is called gradient checking, and the central difference is what you use for it.

The other half: adding up slices

area.py
import numpy as np

def bell(x):
    return np.exp(-x * x / 2) / np.sqrt(2 * np.pi)   # the standard normal curve

for n in (4, 40, 4000):
    edges = np.linspace(-1.0, 1.0, n + 1)
    width = edges[1] - edges[0]
    middles = (edges[:-1] + edges[1:]) / 2
    area = float((bell(middles) * width).sum())
    print(f"{n:>5} rectangles: area {area:.6f}")
print("true value        : 0.682689")
Output
    4 rectangles: area 0.687806
   40 rectangles: area 0.682740
 4000 rectangles: area 0.682689
true value        : 0.682689

That area is the familiar "68% of values fall within one standard deviation". It is a probability, and a probability over a continuous range is an area under a curve.

Notice how fast it converges. Four rectangles get three digits right. This is the midpoint rule, and its error shrinks with the square of the slice width.

The same summing shows up as expectations, as the normalising constant in Bayes' theorem, and inside diffusion models. Unlike the slope half, it is usually approximated by sampling rather than by rectangles. The dimension is far too high for a grid.

The corner in ReLU

corner.py
def relu(x):
    return max(0.0, x)

h = 1e-6
print(f"{'x':>6} {'slope from left':>18} {'slope from right':>18}")
for x0 in (-2.0, 2.0, 0.0):
    left = (relu(x0) - relu(x0 - h)) / h
    right = (relu(x0 + h) - relu(x0)) / h
    print(f"{x0:>6.1f} {left:>18.1f} {right:>18.1f}")
Output
     x    slope from left   slope from right
  -2.0                0.0                0.0
   2.0                1.0                1.0
   0.0                0.0                1.0

At -2.0 and at 2.0 both sides agree, so the slope exists and is well defined.

At exactly 0.0 they disagree: 0.0 from the left, 1.0 from the right. There is no single slope there. ReLU is not differentiable at zero, and this output is the proof rather than an assertion.

Frameworks handle it by picking one. PyTorch returns 0 for the gradient of relu at exactly zero. That is a subgradient — a valid choice from the range of slopes at a corner, not a derivative.

Why it causes no trouble in practice: hitting exactly 0.0 in float32 is vanishingly rare. The choice at that one point does not change the direction of learning.

This is untidy mathematics that works. Saying so is more honest than pretending the corner is not there. Full detail in activation functions.

Common mistakes

1. Choosing h too small in a gradient check. More precision seems safer and is not, as the table above shows. Use h = 1e-5 with a central difference, and compare with a relative tolerance rather than an absolute one.

2. Comparing gradients with ==. Floating-point results almost never match exactly. Use np.allclose(analytic, numeric, rtol=1e-4).

3. Doing a gradient check in float32. There is not enough precision for the subtraction to survive. Cast to float64 for the check, then cast back.

4. Expecting a smooth loss curve. Loss over training steps is noisy, because each step sees a different batch. The underlying function is smooth; your measurement of it is not.

5. Assuming a flat gradient means you are at the bottom. A flat spot can be a minimum, a maximum, or a saddle. In high dimensions saddles are far more common — see optimization.

Try it yourself

Change f in slope.py to x * x * x and predict the true slope at x = 3 before running. Check whether the central difference is still exact at h = 1e-1, and work out why it no longer is.

Then in area.py, change the range from -1.0, 1.0 to -2.0, 2.0. Predict roughly what fraction of the bell curve you will capture, then run it.

What to learn next

Researcher — Mathematics and papers.

Limits and differentiability

The derivative of f at a is defined as a limit, when it exists:

f'(a)  =  lim_{h -> 0}  ( f(a + h) - f(a) ) / h

Existence requires the left and right limits to agree, which is exactly what the ReLU experiment above tests. Differentiability implies continuity; the converse fails. The Weierstrass function is continuous everywhere and differentiable nowhere.

For multivariate f: R**n -> R, the partial derivative varies one coordinate. Existence of all partials does not imply differentiability. The correct notion is Fréchet differentiability: there exists a linear map Df(a) with

lim_{||v|| -> 0}   ( f(a + v) - f(a) - Df(a) v ) / ||v||   =   0

A sufficient condition worth remembering: if all partials exist and are continuous in a neighbourhood, f is differentiable there.

Taylor expansion, the workhorse

Every local analysis in optimisation is a truncated Taylor series. For f: R**n -> R twice differentiable:

f(x + d)  =  f(x)  +  grad_f(x).T @ d  +  (1/2) d.T @ H(x) @ d  +  o(||d||**2)
  • grad_f(x) is the gradient, the vector of first partials.
  • H(x) is the Hessian, the matrix of second partials, symmetric when f is C**2 (Clairaut's theorem).
  • d is a displacement from x.

The first-order term justifies gradient descent: moving along -grad_f decreases f for a small enough step. The second-order term is what Newton's method uses, and what sets the largest usable learning rate.

Smoothness, and where the learning-rate bound comes from

f is L-smooth when its gradient is Lipschitz continuous:

|| grad_f(x) - grad_f(y) ||  <=  L * || x - y ||     for all x, y

Equivalently, when f is C**2, L bounds the largest eigenvalue of the Hessian in absolute value. L-smoothness gives the descent lemma:

f(y)  <=  f(x)  +  grad_f(x).T @ (y - x)  +  (L/2) * || y - x ||**2

Substituting the gradient step y = x - eta * grad_f(x) gives

f(y)  <=  f(x)  -  eta ( 1 - eta L / 2 ) * || grad_f(x) ||**2

The bracket is positive exactly when eta < 2/L, and the decrease is largest at eta = 1/L. That is the origin of the divergence threshold you can watch in practice. For a quadratic with Hessian H, L = lambda_max(H). So the largest curvature direction sets the stable learning rate. The smallest, lambda_min, sets how fast you make progress. Their ratio is the condition number, and it governs the convergence rate.

Deep networks are not globally L-smooth. Cohen et al. (2021) documented the edge of stability. Gradient descent drives lambda_max of the Hessian up to roughly 2/eta, then hovers there. Classical theory does not cover that regime, and the loss decreases non-monotonically inside it.

Convexity

f is convex when f(t x + (1-t) y) <= t f(x) + (1-t) f(y) for t in [0,1]. For differentiable f, this is equivalent to f(y) >= f(x) + grad_f(x).T (y - x): the function lies above every tangent plane. For C**2 functions it is equivalent to a positive semi-definite Hessian everywhere.

Convexity guarantees every local minimum is global. Linear and logistic regression with convex regularisers are convex. Neural networks are not, and no amount of wishing makes them so.

Jensen's inequality — f(E[X]) <= E[f(X)] for convex f — is used constantly. It is the step that produces the ELBO from the log marginal likelihood. It is also why the log-sum-exp trick bounds what it bounds.

Non-smooth points, treated properly

For convex f, the subdifferential at x is the set of g with f(y) >= f(x) + g.T (y - x) for all y. At a differentiable point it is the single gradient. For ReLU at zero it is [0, 1]; frameworks pick 0 by convention.

Two justifications are usually offered, and they differ in strength. The weak one is that the non-differentiable set has Lebesgue measure zero, so it is almost surely never hit. The stronger one is Clarke's generalised gradient theory. Bolte and Pauwels (2020) extended it to the composite functions used in autodiff. They showed backpropagation returns a Clarke subgradient almost everywhere, for a broad class of programs. Kakade and Lee (2018) give a randomised smoothing method with guarantees where this fails.

L1 regularisation raises the same issue at zero, and there the answer is different. Proximal methods handle the kink exactly. That is why they produce genuine sparsity, where subgradient descent leaves small non-zero values.

The integral half

For a continuous random variable with density p, expectation is an integral:

E[g(X)]  =  integral over R**n of  g(x) p(x) dx

Deterministic quadrature costs O(m**n) grid points for m per dimension, so it is unusable above a handful of dimensions. Monte Carlo estimation has error O(1/sqrt(N)) independent of n, which is why every high-dimensional expectation in machine learning is sampled.

Two calculus results carry more weight in practice than they are usually given credit for.

Change of variables. For an invertible, differentiable T with y = T(x):

p_Y(y)  =  p_X( T_inverse(y) ) * | det( d T_inverse / d y ) |

This is the whole basis of normalising flows. Architectural choices there are driven entirely by making that Jacobian determinant cheap to compute.

Differentiating under the integral sign. Swapping d/d theta and integral is licensed by the dominated convergence theorem under mild conditions. When the density itself depends on theta, the swap is invalid as written. The reparameterisation trick (Kingma and Welling, 2013) restores it. It pushes the parameter dependence into a deterministic transform of parameter-free noise. The alternative is the score-function or REINFORCE estimator: unbiased, and far higher variance.

References

  • Spivak, M. Calculus, 4th ed., Publish or Perish, 2008. Rigorous and readable.
  • Rudin, W. Principles of Mathematical Analysis, 3rd ed., McGraw-Hill, 1976. Chapters 5 and 9.
  • Nesterov, Y. Lectures on Convex Optimization, 2nd ed., Springer, 2018. The descent lemma and its consequences.
  • Boyd, S., Vandenberghe, L. Convex Optimization. Cambridge University Press, 2004.
  • Clarke, F. H. Optimization and Nonsmooth Analysis. Wiley, 1983.
  • Bolte, J., Pauwels, E. "A mathematical model for automatic differentiation in machine learning." NeurIPS, 2020. arXiv:2006.02080
  • Kakade, S. M., Lee, J. D. "Provably Correct Automatic Subdifferentiation for Qualified Programs." NeurIPS, 2018. arXiv:1809.08530
  • Cohen, J. M., Kaur, S., Li, Y., Kolter, J. Z., Talwalkar, A. "Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability." ICLR, 2021. arXiv:2103.00065
  • Kingma, D. P., Welling, M. "Auto-Encoding Variational Bayes." arXiv:1312.6114, 2013.

What to learn next