Why AI needs mathematics
You do not need to be a mathematician to build AI. You need four small ideas, mainly so you can tell why a model failed instead of guessing.
- 13 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
You do not need to be a mathematician to build AI. You need four small ideas, and mostly so you can work out why something went wrong.
Think about learning to drive. You can reach the market without knowing what happens under the bonnet. But one day on the highway, the car makes a strange grinding noise. The driver who understands the engine gets home. The one who does not waits at the roadside.
AI libraries are the car. They do the mathematics for you, and they do it well. You will get surprisingly far without opening the bonnet.
Why you should care
Here is the part nobody warns you about. AI models rarely crash when they are wrong.
A broken web page shows an error. A broken model gives you a clean, confident, wrong answer, in a nicely formatted table, and says nothing. It looks exactly like a working model.
The mathematics is what lets you tell the two apart. Not to build the engine — to hear the noise.
Why it exists
Every AI model does one loop, over and over:
Make a guess → see how wrong it was → change slightly to be less wrong → repeatEach arrow in that loop needed a precise answer to a hard question. How do we hold millions of measurements and combine them fast? How do we measure wrongness in a single number? How do we know which direction is "less wrong"? How much should we change?
Those questions already had answers, worked out over three hundred years by people studying planets, gambling, and heat. AI did not invent new mathematics so much as it found a new use for old mathematics.
The four things you actually need
LINEAR ALGEBRA → how to hold huge piles of measurements and combine them at once
CALCULUS → how to work out which way to change to become less wrong
PROBABILITY → how to say "I am not fully sure" instead of pretending
STATISTICS → how to check whether a result is real or a flukeEvery one of them shows up in the loop above:
Your data → becomes rows of measurements (linear algebra)
The model → combines those rows into a guess (linear algebra)
The guess → gets compared with the true answer (statistics)
The mistake → points toward which way to adjust (calculus)
The adjusting → happens in small steps, many times (all of them)A real example you have seen
Your bank sends a message: "we noticed unusual activity on your card". Notice the wording. Not "this is fraud" — "unusual".
The system did not know. It worked out that this purchase was unlike your usual ones, and that being cautious costs less than being wrong. Somebody chose how suspicious is suspicious enough. That choice is probability, and it is why your card gets blocked on holiday sometimes.
An honest word
You do not need all four before you start. Waiting until you "know the maths" is the most common way people never begin.
Start building. Come back here when something breaks. These ideas stick far better attached to a real failure than read once as a chapter.
Also worth saying plainly: school mathematics is taught as a set of procedures to be performed under time pressure. This is not that. Here you need to understand what four ideas mean, and a computer does the arithmetic. Many people who disliked school mathematics find they are fine with this.
Remember this
- Libraries do the mathematics for you, so you rarely calculate anything by hand.
- You need the ideas because a wrong model looks exactly like a right one.
- Four areas cover almost everything: measurements, direction of change, uncertainty, and evidence.
What to learn next
- Linear algebra — the language of many measurements at once.
- Vectors and matrices — the two objects everything is built from.
- Statistics — telling a real result from a lucky one.
Developer — Code and libraries.
Setup
pip install numpyThe point of this block is to make the four areas concrete. Below is a complete model that learns, written from scratch with no machine-learning library. It runs on a CPU in well under a second.
Every branch of the maths, in twenty lines
import numpy as np
# Six flats. size = hundreds of square feet, price = lakhs of rupees.
size = np.array([ 5.0, 7.0, 8.0, 10.0, 12.0, 15.0])
price = np.array([25.0, 34.0, 41.0, 48.0, 60.0, 73.0])
w, b = 0.0, 0.0 # the two numbers the model must learn
lr = 0.005 # learning rate: how big a step we take each round
for step in range(1, 2001):
pred = w * size + b
error = pred - price
loss = (error ** 2).mean() # this is statistics
dw = 2.0 * (error * size).mean() # these two lines are calculus
db = 2.0 * error.mean()
w -= lr * dw # this is optimization
b -= lr * db
if step in (1, 10, 100, 500, 2000):
print(f"step {step:>4} loss {loss:9.4f} w {w:6.3f} b {b:6.3f}")
print()
print(f"learned rule : price = {w:.3f} * size + {b:.3f}")
slope, intercept = np.polyfit(size, price, 1)
print(f"exact answer : price = {slope:.3f} * size + {intercept:.3f}")step 1 loss 2449.1667 w 4.977 b 0.468 step 10 loss 0.9378 w 4.876 b 0.463 step 100 loss 0.9336 w 4.872 b 0.506 step 500 loss 0.9225 w 4.858 b 0.653 step 2000 loss 0.9147 w 4.837 b 0.873 learned rule : price = 4.837 * size + 0.873 exact answer : price = 4.832 * size + 0.929
That is a complete machine-learning system. It saw six examples and worked out a rule. That rule nearly matches the exact least-squares answer, which np.polyfit computes in closed form.
Line-by-line walkthrough
pred = w * size + b — linear algebra. One line produces six predictions because size is an array. In a real network w is a matrix and size is a batch of examples. This line then becomes a matrix multiply. The idea does not change; the shapes do.
loss = (error ** 2).mean() — statistics. This turns six separate mistakes into one number to minimise. Squaring does two things. It stops overshooting and undershooting from cancelling out. It also punishes one large mistake far more than several small ones. That choice has consequences — squared error is sensitive to outliers, which is why other loss functions exist.
dw and db — calculus. These are the derivatives of the loss with respect to w and b. A derivative answers one question. Nudge this number up a little: does the loss rise or fall, and how sharply? The sign tells you the direction to move. The size tells you how confident that direction is.
w -= lr * dw — optimization. Move against the gradient, because the gradient points uphill and you want down. The learning rate lr decides how far to step.
Watch the loss column. It drops from 2449 to 0.94 between step 1 and step 10, then barely moves for the next 1990 steps. That plateau is normal and it is not the model being lazy. Near a minimum the gradient is small, so the steps are small. Most of the improvement in most training runs happens early.
Common mistakes
1. Learning rate too large.
Change lr to 0.02 in the code above and the same model explodes:
step 1 loss 2.449e+03 w 1.991e+01 step 5 loss 1.995e+07 w 1.362e+03 step 10 loss 1.545e+12 w -3.775e+05 step 20 loss 9.259e+21 w -2.923e+10
Each step overshoots the bottom and lands further up the other side. Eventually you get inf or nan. A loss that grows instead of shrinking is almost always this.
2. Features on wildly different scales. Say one input is in rupees (tens of thousands) and another is a count (zero to five). The gradient for the first dwarfs the second. The model chases the big one and ignores the small one. Fix: subtract the mean and divide by the standard deviation for each column before training.
3. Judging a model by its training loss. The loss above says how well the model memorised six flats. It says nothing about a seventh flat. Every honest evaluation needs data the model has not seen.
4. Reaching for calculus by hand.
You worked out dw and db here because the model has two parameters. A real network has millions, and nobody differentiates those by hand. Autograd in PyTorch or JAX does it automatically. Deriving it once, as above, is worth doing so the automatic version stops being magic.
Try it yourself
Add a seventh flat with a size of 20.0 and a price of 130.0. It is an outlier, priced far above the trend. Rerun and watch how much the learned slope shifts for one bad row. That single experiment teaches more about squared error than a page of text.
Then replace (error ** 2).mean() with np.abs(error).mean(), and change the gradients to (np.sign(error) * size).mean() and np.sign(error).mean(). Compare how much the outlier moves the answer now.
What to learn next
- Vectors and matrices — the shapes that
w * sizebecomes at scale. - Derivatives and gradients — where
dwcomes from. - Gradient descent — the update rule in full.
Researcher — Mathematics and papers.
The problem, stated once
Supervised learning is empirical risk minimisation. Samples are drawn independently from an unknown joint distribution over inputs and labels. Choose the parameters that minimise average loss on those samples:
R_emp(theta) = (1/n) * sum_{i=1..n} L( f(x_i; theta), y_i )nis the number of training examples.x_iis thei-th input,y_iits true label.thetais the full parameter vector of the model.f(x; theta)is the model's prediction for inputx.L(a, b)is the loss, a non-negative function measuring how bad predictionais against truthb.
The quantity actually wanted is the true risk, the expected loss over the whole distribution D:
R(theta) = E_{(x,y) ~ D} [ L( f(x; theta), y ) ]Everything difficult in machine learning lives in the gap R(theta) - R_emp(theta). Statistics is the study of that gap.
Where each branch enters
Linear algebra supplies the model class and the cost model. A dense layer maps a batch X of shape (B, d_in) to
Z = X @ W + b W has shape (d_in, d_out), b has shape (d_out,)with a forward cost of about 2 * B * d_in * d_out floating-point operations, counting a multiply-accumulate as two. Backward is roughly twice forward, so a training step costs about three times a forward pass. This is the arithmetic behind every FLOP budget you will read in a model card.
Calculus supplies the descent direction. Gradient descent is:
theta_{t+1} = theta_t - eta * grad_theta R_emp(theta_t)etais the learning rate, a positive scalar.grad_theta R_empis the vector of partial derivatives of the empirical risk with respect to every entry oftheta.
Take an L-smooth convex objective, meaning its gradient does not change faster than a constant L. The classical condition for monotone decrease is then eta < 2/L. The divergence shown in the Developer block is exactly this bound being violated.
Reverse-mode automatic differentiation computes the full gradient within a small constant factor of one forward pass. That single result is what makes deep learning computationally possible.
Probability supplies the loss functions, which are not arbitrary. Maximum likelihood estimation chooses
theta_MLE = argmax_theta sum_{i=1..n} log p( y_i | x_i ; theta )
= argmin_theta -(1/n) * sum_{i=1..n} log p( y_i | x_i ; theta )where p(y | x; theta) is the model's assumed conditional density or mass function. Two familiar losses fall straight out:
- Assume
y | xis Gaussian with meanf(x; theta)and fixed variance. The negative log-likelihood reduces to mean squared error plus a constant. - Assume
y | xis categorical with class probabilities from a softmax. The negative log-likelihood is cross-entropy.
So "use MSE for regression, cross-entropy for classification" is not a convention. It is a statement about which noise model you are assuming.
Bayes' theorem gives the machinery for updating belief with evidence:
P(A | B) = P(B | A) * P(A) / P(B)P(A)is the prior, belief before seeingB.P(B | A)is the likelihood of the evidence underA.P(A | B)is the posterior, belief after seeingB.
This underlies naive Bayes, Bayesian optimisation for hyperparameter search, Thompson sampling, and variational inference.
Statistics supplies the guarantee that any of this transfers. With a hypothesis class of finite VC dimension d and n samples, uniform convergence gives a bound of the form
R(theta) <= R_emp(theta) + O( sqrt( (d + log(1/delta)) / n ) )holding with probability at least 1 - delta. The bound is loose for modern overparameterised networks. Explaining why they generalise anyway remains an open research area; see Zhang et al., 2017.
The qualitative lesson survives. Generalisation improves with more data and degrades with more model capacity. And any claim about a model needs a confidence interval, not a point estimate.
The minimum you should be fluent in
| Idea | Why it appears | Typical first encounter |
|---|---|---|
| Matrix multiplication, shapes, transpose | every layer, every attention head | dense layers |
| Norms and inner products | similarity, regularisation, gradient clipping | embeddings, cosine similarity |
| Eigendecomposition and SVD | PCA, low-rank adapters, conditioning | LoRA, whitening |
| Partial derivatives and the chain rule | backpropagation | training any network |
| Convexity and smoothness | why a learning rate works or does not | optimiser choice |
| Expectation, variance, covariance | batch statistics, normalisation layers | BatchNorm, LayerNorm |
| Conditional probability, Bayes | losses, calibration, uncertainty | cross-entropy, temperature |
| Sampling and confidence intervals | evaluation you can defend | benchmark reporting |
References
- Deisenroth, M. P., Faisal, A. A., Ong, C. S. Mathematics for Machine Learning. Cambridge University Press, 2020. Freely available from the authors, and the best single starting point.
- Goodfellow, I., Bengio, Y., Courville, A. Deep Learning. MIT Press, 2016. Chapters 2 to 4 are a compact survey of exactly this material.
- Shalev-Shwartz, S., Ben-David, S. Understanding Machine Learning: From Theory to Algorithms. Cambridge University Press, 2014. The generalisation bounds, done properly.
- Boyd, S., Vandenberghe, L. Convex Optimization. Cambridge University Press, 2004.
- Robbins, H., Monro, S. "A Stochastic Approximation Method." Annals of Mathematical Statistics 22(3), 400–407, 1951. The origin of stochastic gradient descent.
- Rumelhart, D. E., Hinton, G. E., Williams, R. J. "Learning representations by back-propagating errors." Nature 323, 533–536, 1986.
- Zhang, C., Bengio, S., Hardt, M., Recht, B., Vinyals, O. "Understanding deep learning requires rethinking generalization." ICLR, 2017.
What to learn next
- Vectors and matrices — the objects and their costs.
- Derivatives and gradients — reverse-mode differentiation in detail.
- Optimization — momentum, Adam, and convergence conditions.