Mathematics for AI

Why AI needs mathematics

You do not need to be a mathematician to build AI. You need four small ideas, mainly so you can tell why a model failed instead of guessing.

Read these first

On this page 7
  1. Why you should care
  2. Why it exists
  3. The four things you actually need
  4. A real example you have seen
  5. An honest word
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

You do not need to be a mathematician to build AI. You need four small ideas, and mostly so you can work out why something went wrong.

Think about learning to drive. You can reach the market without knowing what happens under the bonnet. But one day on the highway, the car makes a strange grinding noise. The driver who understands the engine gets home. The one who does not waits at the roadside.

AI libraries are the car. They do the mathematics for you, and they do it well. You will get surprisingly far without opening the bonnet.

Why you should care

Here is the part nobody warns you about. AI models rarely crash when they are wrong.

A broken web page shows an error. A broken model gives you a clean, confident, wrong answer, in a nicely formatted table, and says nothing. It looks exactly like a working model.

The mathematics is what lets you tell the two apart. Not to build the engine — to hear the noise.

Why it exists

Every AI model does one loop, over and over:

Make a guess  →  see how wrong it was  →  change slightly to be less wrong  →  repeat

Each arrow in that loop needed a precise answer to a hard question. How do we hold millions of measurements and combine them fast? How do we measure wrongness in a single number? How do we know which direction is "less wrong"? How much should we change?

Those questions already had answers, worked out over three hundred years by people studying planets, gambling, and heat. AI did not invent new mathematics so much as it found a new use for old mathematics.

The four things you actually need

LINEAR ALGEBRA  →  how to hold huge piles of measurements and combine them at once
CALCULUS        →  how to work out which way to change to become less wrong
PROBABILITY     →  how to say "I am not fully sure" instead of pretending
STATISTICS      →  how to check whether a result is real or a fluke

Every one of them shows up in the loop above:

Your data     →  becomes rows of measurements               (linear algebra)
The model     →  combines those rows into a guess           (linear algebra)
The guess     →  gets compared with the true answer         (statistics)
The mistake   →  points toward which way to adjust          (calculus)
The adjusting →  happens in small steps, many times         (all of them)

A real example you have seen

Your bank sends a message: "we noticed unusual activity on your card". Notice the wording. Not "this is fraud" — "unusual".

The system did not know. It worked out that this purchase was unlike your usual ones, and that being cautious costs less than being wrong. Somebody chose how suspicious is suspicious enough. That choice is probability, and it is why your card gets blocked on holiday sometimes.

An honest word

You do not need all four before you start. Waiting until you "know the maths" is the most common way people never begin.

Start building. Come back here when something breaks. These ideas stick far better attached to a real failure than read once as a chapter.

Also worth saying plainly: school mathematics is taught as a set of procedures to be performed under time pressure. This is not that. Here you need to understand what four ideas mean, and a computer does the arithmetic. Many people who disliked school mathematics find they are fine with this.

Remember this

  • Libraries do the mathematics for you, so you rarely calculate anything by hand.
  • You need the ideas because a wrong model looks exactly like a right one.
  • Four areas cover almost everything: measurements, direction of change, uncertainty, and evidence.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

The point of this block is to make the four areas concrete. Below is a complete model that learns, written from scratch with no machine-learning library. It runs on a CPU in well under a second.

Every branch of the maths, in twenty lines

four_ideas.py
import numpy as np

# Six flats. size = hundreds of square feet, price = lakhs of rupees.
size  = np.array([ 5.0,  7.0,  8.0, 10.0, 12.0, 15.0])
price = np.array([25.0, 34.0, 41.0, 48.0, 60.0, 73.0])

w, b = 0.0, 0.0     # the two numbers the model must learn
lr = 0.005          # learning rate: how big a step we take each round

for step in range(1, 2001):
    pred = w * size + b
    error = pred - price
    loss = (error ** 2).mean()                 # this is statistics
    dw = 2.0 * (error * size).mean()           # these two lines are calculus
    db = 2.0 * error.mean()
    w -= lr * dw                               # this is optimization
    b -= lr * db
    if step in (1, 10, 100, 500, 2000):
        print(f"step {step:>4}  loss {loss:9.4f}  w {w:6.3f}  b {b:6.3f}")

print()
print(f"learned rule : price = {w:.3f} * size + {b:.3f}")
slope, intercept = np.polyfit(size, price, 1)
print(f"exact answer : price = {slope:.3f} * size + {intercept:.3f}")
Output
step    1  loss 2449.1667  w  4.977  b  0.468
step   10  loss    0.9378  w  4.876  b  0.463
step  100  loss    0.9336  w  4.872  b  0.506
step  500  loss    0.9225  w  4.858  b  0.653
step 2000  loss    0.9147  w  4.837  b  0.873

learned rule : price = 4.837 * size + 0.873
exact answer : price = 4.832 * size + 0.929

That is a complete machine-learning system. It saw six examples and worked out a rule. That rule nearly matches the exact least-squares answer, which np.polyfit computes in closed form.

Line-by-line walkthrough

pred = w * size + b — linear algebra. One line produces six predictions because size is an array. In a real network w is a matrix and size is a batch of examples. This line then becomes a matrix multiply. The idea does not change; the shapes do.

loss = (error ** 2).mean() — statistics. This turns six separate mistakes into one number to minimise. Squaring does two things. It stops overshooting and undershooting from cancelling out. It also punishes one large mistake far more than several small ones. That choice has consequences — squared error is sensitive to outliers, which is why other loss functions exist.

dw and db — calculus. These are the derivatives of the loss with respect to w and b. A derivative answers one question. Nudge this number up a little: does the loss rise or fall, and how sharply? The sign tells you the direction to move. The size tells you how confident that direction is.

w -= lr * dw — optimization. Move against the gradient, because the gradient points uphill and you want down. The learning rate lr decides how far to step.

Watch the loss column. It drops from 2449 to 0.94 between step 1 and step 10, then barely moves for the next 1990 steps. That plateau is normal and it is not the model being lazy. Near a minimum the gradient is small, so the steps are small. Most of the improvement in most training runs happens early.

Common mistakes

1. Learning rate too large. Change lr to 0.02 in the code above and the same model explodes:

Output
step  1  loss 2.449e+03  w 1.991e+01
step  5  loss 1.995e+07  w 1.362e+03
step 10  loss 1.545e+12  w -3.775e+05
step 20  loss 9.259e+21  w -2.923e+10

Each step overshoots the bottom and lands further up the other side. Eventually you get inf or nan. A loss that grows instead of shrinking is almost always this.

2. Features on wildly different scales. Say one input is in rupees (tens of thousands) and another is a count (zero to five). The gradient for the first dwarfs the second. The model chases the big one and ignores the small one. Fix: subtract the mean and divide by the standard deviation for each column before training.

3. Judging a model by its training loss. The loss above says how well the model memorised six flats. It says nothing about a seventh flat. Every honest evaluation needs data the model has not seen.

4. Reaching for calculus by hand. You worked out dw and db here because the model has two parameters. A real network has millions, and nobody differentiates those by hand. Autograd in PyTorch or JAX does it automatically. Deriving it once, as above, is worth doing so the automatic version stops being magic.

Try it yourself

Add a seventh flat with a size of 20.0 and a price of 130.0. It is an outlier, priced far above the trend. Rerun and watch how much the learned slope shifts for one bad row. That single experiment teaches more about squared error than a page of text.

Then replace (error ** 2).mean() with np.abs(error).mean(), and change the gradients to (np.sign(error) * size).mean() and np.sign(error).mean(). Compare how much the outlier moves the answer now.

What to learn next

Researcher — Mathematics and papers.

The problem, stated once

Supervised learning is empirical risk minimisation. Samples are drawn independently from an unknown joint distribution over inputs and labels. Choose the parameters that minimise average loss on those samples:

R_emp(theta)  =  (1/n) * sum_{i=1..n}  L( f(x_i; theta),  y_i )
  • n is the number of training examples.
  • x_i is the i-th input, y_i its true label.
  • theta is the full parameter vector of the model.
  • f(x; theta) is the model's prediction for input x.
  • L(a, b) is the loss, a non-negative function measuring how bad prediction a is against truth b.

The quantity actually wanted is the true risk, the expected loss over the whole distribution D:

R(theta)  =  E_{(x,y) ~ D} [ L( f(x; theta), y ) ]

Everything difficult in machine learning lives in the gap R(theta) - R_emp(theta). Statistics is the study of that gap.

Where each branch enters

Linear algebra supplies the model class and the cost model. A dense layer maps a batch X of shape (B, d_in) to

Z  =  X @ W + b          W has shape (d_in, d_out),  b has shape (d_out,)

with a forward cost of about 2 * B * d_in * d_out floating-point operations, counting a multiply-accumulate as two. Backward is roughly twice forward, so a training step costs about three times a forward pass. This is the arithmetic behind every FLOP budget you will read in a model card.

Calculus supplies the descent direction. Gradient descent is:

theta_{t+1}  =  theta_t  -  eta * grad_theta R_emp(theta_t)
  • eta is the learning rate, a positive scalar.
  • grad_theta R_emp is the vector of partial derivatives of the empirical risk with respect to every entry of theta.

Take an L-smooth convex objective, meaning its gradient does not change faster than a constant L. The classical condition for monotone decrease is then eta < 2/L. The divergence shown in the Developer block is exactly this bound being violated.

Reverse-mode automatic differentiation computes the full gradient within a small constant factor of one forward pass. That single result is what makes deep learning computationally possible.

Probability supplies the loss functions, which are not arbitrary. Maximum likelihood estimation chooses

theta_MLE  =  argmax_theta  sum_{i=1..n}  log p( y_i | x_i ; theta )
           =  argmin_theta  -(1/n) * sum_{i=1..n}  log p( y_i | x_i ; theta )

where p(y | x; theta) is the model's assumed conditional density or mass function. Two familiar losses fall straight out:

  • Assume y | x is Gaussian with mean f(x; theta) and fixed variance. The negative log-likelihood reduces to mean squared error plus a constant.
  • Assume y | x is categorical with class probabilities from a softmax. The negative log-likelihood is cross-entropy.

So "use MSE for regression, cross-entropy for classification" is not a convention. It is a statement about which noise model you are assuming.

Bayes' theorem gives the machinery for updating belief with evidence:

P(A | B)  =  P(B | A) * P(A) / P(B)
  • P(A) is the prior, belief before seeing B.
  • P(B | A) is the likelihood of the evidence under A.
  • P(A | B) is the posterior, belief after seeing B.

This underlies naive Bayes, Bayesian optimisation for hyperparameter search, Thompson sampling, and variational inference.

Statistics supplies the guarantee that any of this transfers. With a hypothesis class of finite VC dimension d and n samples, uniform convergence gives a bound of the form

R(theta)  <=  R_emp(theta)  +  O( sqrt( (d + log(1/delta)) / n ) )

holding with probability at least 1 - delta. The bound is loose for modern overparameterised networks. Explaining why they generalise anyway remains an open research area; see Zhang et al., 2017.

The qualitative lesson survives. Generalisation improves with more data and degrades with more model capacity. And any claim about a model needs a confidence interval, not a point estimate.

The minimum you should be fluent in

IdeaWhy it appearsTypical first encounter
Matrix multiplication, shapes, transposeevery layer, every attention headdense layers
Norms and inner productssimilarity, regularisation, gradient clippingembeddings, cosine similarity
Eigendecomposition and SVDPCA, low-rank adapters, conditioningLoRA, whitening
Partial derivatives and the chain rulebackpropagationtraining any network
Convexity and smoothnesswhy a learning rate works or does notoptimiser choice
Expectation, variance, covariancebatch statistics, normalisation layersBatchNorm, LayerNorm
Conditional probability, Bayeslosses, calibration, uncertaintycross-entropy, temperature
Sampling and confidence intervalsevaluation you can defendbenchmark reporting

References

  • Deisenroth, M. P., Faisal, A. A., Ong, C. S. Mathematics for Machine Learning. Cambridge University Press, 2020. Freely available from the authors, and the best single starting point.
  • Goodfellow, I., Bengio, Y., Courville, A. Deep Learning. MIT Press, 2016. Chapters 2 to 4 are a compact survey of exactly this material.
  • Shalev-Shwartz, S., Ben-David, S. Understanding Machine Learning: From Theory to Algorithms. Cambridge University Press, 2014. The generalisation bounds, done properly.
  • Boyd, S., Vandenberghe, L. Convex Optimization. Cambridge University Press, 2004.
  • Robbins, H., Monro, S. "A Stochastic Approximation Method." Annals of Mathematical Statistics 22(3), 400–407, 1951. The origin of stochastic gradient descent.
  • Rumelhart, D. E., Hinton, G. E., Williams, R. J. "Learning representations by back-propagating errors." Nature 323, 533–536, 1986.
  • Zhang, C., Bengio, S., Hardt, M., Recht, B., Vinyals, O. "Understanding deep learning requires rethinking generalization." ICLR, 2017.

What to learn next

What to learn next

These follow on from what you just read.

  • Mathematics for AI

    Linear algebra

    Linear algebra is the maths of scaling things and adding them up. Every layer of every AI model is that one move, repeated at enormous scale.

  • Mathematics for AI

    Vectors and matrices

    A vector is an ordered list of measurements about one thing. A matrix is many of those lists stacked into a table. Almost everything inside an AI model is one of these two.

  • Mathematics for AI

    Probability

    Probability is how a model says how sure it is. Every confidence score, every spam filter and every fraud alert is built on it.