Mathematics for AI

Probability

Probability is how a model says how sure it is. Every confidence score, every spam filter and every fraud alert is built on it.

On this page 8
  1. Why you should care
  2. The scale
  3. Why it exists
  4. The trap everybody falls into
  5. Where you have already used it
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Probability is a number that says how sure you are, on one fixed scale.

Think about stepping outside in the morning and looking up at the sky. Grey clouds, heavy air. You go back in and pick up the umbrella.

You did not know it would rain. You weighed the signs and acted on a hunch. Probability is that hunch, written down as a number everyone reads the same way.

Why you should care

An AI model never gives you an answer. It gives you a confidence, and something downstream turns that confidence into an answer.

The spam filter does not know an email is spam. It leans that way. Somebody picked a cut-off, and above that cut-off the email goes to the spam folder.

Once you see that, a whole class of AI failures becomes readable. A model that is right most of the time can still be badly wrong about how sure it is.

The scale

Every probability sits between two ends. One end means it never happens. The other end means it always happens.

   never                      even chance                     always
     |-------------------------------|-------------------------------|
   the sun               a fair coin landing               the sun
   rising in                    heads                       rising in
   the west                                                 the east

Nothing sits outside those two ends. That is the whole scale, and every confidence score you will ever see lives on it.

There is a second rule and it is nearly as short. List every possible outcome, with no overlaps and nothing missing. Their chances add up to the "always" end of the scale, exactly.

Why it exists

Probability was worked out for an unglamorous reason. Gamblers in seventeenth-century France kept arguing about how to split the stake when a game was stopped halfway.

Two mathematicians, Blaise Pascal and Pierre de Fermat, wrote to each other about it. Out of those letters came a way to reason about partial information.

That is still the job. You almost never have all the facts. Probability is the discipline of acting well anyway.

The trap everybody falls into

This next part is confusing for almost everyone the first time. Read it twice; that is normal, not a sign you are slow.

Suppose a fraud alarm is very good. It rings for nearly every real fraud. It also rings, rarely, for honest payments — say a couple out of every hundred.

Your card gets blocked. The alarm has rung. How likely is it that you are actually being defrauded?

Most people answer "very likely". The truth is the opposite.

   Out of a hundred thousand payments in a day:

   real fraud    about a hundred      →  the alarm catches nearly all
   honest        almost all the rest  →  the alarm wrongly rings for
                                         a couple out of every hundred

   so today's alarms look roughly like this:

   from real fraud   #
   from honest       ####################

   nearly every alarm comes from an honest payment

Honest payments are so much more common that even a rare mistake on them swamps the real cases. Fewer than one alarm in twenty is genuine.

Nothing about the alarm is broken. The rarity of fraud is doing all the work.

That rarity has a name. It is the base rate: how common a thing is before you see any evidence.

This single idea explains medical screening, spam filters and fraud alerts. Ignore the base rate and you will trust a rare-disease test far more than you should.

Where you have already used it

  • Your weather app, which reports a chance of rain, not a promise.
  • The spam folder, which is a cut-off applied to a confidence score.
  • Autocorrect, ranking which word you probably meant.
  • A bank message saying "unusual activity", which is careful wording chosen because the bank knows it is often wrong.

The honest part

A model's confidence is often wrong, and wrong in a predictable direction. Large models tend to be overconfident. Something reported as near-certain turns out to be right rather less often than that.

There is a name for a model whose confidences can be trusted at face value: calibrated. Most models are not, out of the box. It can be fixed, and the fix is measured, not assumed.

So treat a confidence score as a rough ranking, not a promise, until somebody has checked it against reality.

Remember this

  • Probability is a number for how sure you are, sitting between "never" and "always".
  • Models output confidences, not answers. A cut-off turns confidence into a decision.
  • How common a thing is beforehand changes everything. Rare things stay rare, even after an alarm.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Three short programs. The first turns raw model output into probabilities. The second is the base-rate trap in code. The third is the reason nobody multiplies probabilities in production.

Raw scores are not probabilities

A classifier's last layer produces logits — unbounded raw scores with no meaning on their own. They can be negative, and they do not add up to anything. Softmax converts them into probabilities.

softmax.py
import numpy as np

def softmax(z, temperature=1.0):
    z = np.asarray(z, dtype=float) / temperature
    z = z - z.max()                 # shift so exp() can never overflow
    e = np.exp(z)
    return e / e.sum()

labels = ["cat", "dog", "rabbit"]
scores = np.array([2.0, 1.0, 0.1])  # raw model outputs, called logits

p = softmax(scores)
print("raw scores    :", scores)
print("probabilities :", np.round(p, 4))
print("they add up to:", round(float(p.sum()), 10))
print("prediction    :", labels[int(p.argmax())])
print()
for t in (0.5, 1.0, 4.0):
    print(f"temperature {t:>4}: {np.round(softmax(scores, t), 4)}")
Output
raw scores    : [2.  1.  0.1]
probabilities : [0.659  0.2424 0.0986]
they add up to: 1.0
prediction    : cat

temperature  0.5: [0.8638 0.1169 0.0193]
temperature  1.0: [0.659  0.2424 0.0986]
temperature  4.0: [0.4165 0.3244 0.259 ]

Two things in that output are worth stopping on.

The z - z.max() line is not decoration. Logits of 1000 would make np.exp return inf, and inf / inf is nan. Subtracting the maximum leaves the answer mathematically identical and removes the overflow. Every production softmax does this.

Temperature changes the confidence, never the ranking. Low temperature sharpens toward the top class. High temperature flattens everything toward equal. cat wins at all three settings. This is exactly the dial described in temperature and sampling.

That second point has a sharp edge. A softmax score is a probability by construction, since it is positive and sums to one. Being a valid probability does not make it a truthful one.

The base-rate trap, in seven lines

base_rate.py
TOTAL = 100_000
fraud = 100                 # about one payment in a thousand is fraud
honest = TOTAL - fraud
caught = 0.99               # the alarm rings for 99% of real fraud
false_alarm = 0.02          # it also rings for 2% of honest payments

rings_fraud = fraud * caught
rings_honest = honest * false_alarm
print("alarms on real fraud   :", round(rings_fraud))
print("alarms on honest people:", round(rings_honest))
print("total alarms           :", round(rings_fraud + rings_honest))
print("chance an alarm is real:", round(rings_fraud / (rings_fraud + rings_honest), 4))
Output
alarms on real fraud   : 99
alarms on honest people: 1998
total alarms           : 2097
chance an alarm is real: 0.0472

A detector that catches 99 out of every 100 frauds produces alarms that are right 4.7% of the time. Nothing is broken. The 2% false-alarm rate is applied to a vastly larger group.

This is Bayes' theorem, worked in whole people instead of fractions. That is the way to teach it, and the way to debug it. The last number is precision: of everything you flagged, what share was real.

Two engineering consequences follow immediately.

Accuracy is a useless metric here. A detector that flags nothing at all is 99.9% accurate on this data. Report precision and recall, and read them together. See model evaluation.

A tiny change in the false-alarm rate moves everything. Set false_alarm = 0.001 and rerun. Precision jumps from 4.7% to about 50%. On rare events, the false-positive rate matters far more than the catch rate.

Never multiply probabilities

The probability of a whole sentence is the product of the probability of each token. Do that directly and it dies.

logprobs.py
import numpy as np

token_probs = np.full(2000, 0.3)     # a 2000-token answer, each token 30% likely
print("multiplied out :", np.prod(token_probs))
print("sum of logs    :", round(float(np.log(token_probs).sum()), 4))
print("perplexity     :", round(float(np.exp(-np.log(token_probs).mean())), 4))
Output
multiplied out : 0.0
sum of logs    : -2407.9456
perplexity     : 3.3333

The direct product underflows to exactly zero. A float64 cannot hold a number that small, so the information is gone. No warning is printed.

Adding logarithms instead keeps every bit of it. This is why model APIs return logprobs and never probs, and why loss functions are written as log-likelihoods.

Perplexity is that log average, exponentiated back. Read it as an effective number of choices. Here every token had a 0.3 chance, so the model was effectively picking between about 3.33 options at each step. Lower is better.

Common mistakes

1. Treating a softmax score as a real-world chance. A model reporting 0.95 may be right 80% of the time. Bin your validation predictions by confidence and compare each bin against how often it was right. That gap is calibration error, and it is measurable in a dozen lines.

2. Fixing the threshold at 0.5. The default is a habit, not a principle. On imbalanced data the useful threshold is set by what a miss costs versus what a false alarm costs.

3. Assuming independence and not saying so. Naive Bayes multiplies word probabilities as though words were unrelated. They are not — "New" and "York" travel together. It still works well for spam, because the ranking survives even when the numbers are wrong.

4. Reading probability into a model that was never trained for it. Support vector machines and unregularised tree ensembles emit scores, not probabilities. Wrap them in a calibrator before you read anything into the value.

5. Averaging probabilities across differently-sized groups. This produces Simpson's paradox, where every subgroup shows one trend and the total shows the opposite.

Try it yourself

Take base_rate.py and turn the fraud rate up to one in ten, leaving everything else alone. Watch precision go from 4.7% to over 84%. The detector did not change. Only the world it runs in did.

Then rewrite softmax without the z - z.max() line and feed it [1000.0, 999.0, 998.0]. You will get nan values and no error message. That is the failure the shift exists to prevent.

What to learn next

Researcher — Mathematics and papers.

Axioms and objects

Kolmogorov's 1933 axioms define a probability space (Omega, F, P). Omega is the sample space and F a sigma-algebra of measurable events. P: F -> [0, 1] is a measure with P(Omega) = 1, countably additive over disjoint events.

A random variable X is a measurable function from Omega to R. It is not a number, and not an unknown. It is a function. Forgetting that causes most of the notational confusion downstream.

Core quantities:

E[X]      =  integral over Omega of X dP           the expectation, or mean
Var[X]    =  E[ (X - E[X])**2 ]  =  E[X**2] - E[X]**2
Cov[X,Y]  =  E[ (X - E[X]) (Y - E[Y]) ]

Linearity of expectation, E[aX + bY] = a E[X] + b E[Y], holds with no independence assumption whatsoever. Variance does not: Var[X + Y] = Var[X] + Var[Y] + 2 Cov[X, Y].

That asymmetry is load-bearing in machine learning. Averaging n correlated predictions cuts variance by less than a factor of n. That is why ensemble diversity is worth engineering. See random forest.

Conditioning and Bayes

P(A | B)  =  P(A and B) / P(B),          for P(B) > 0

P(A | B)  =  P(B | A) P(A) / P(B)
          =  P(B | A) P(A) / ( P(B|A) P(A) + P(B|not A) P(not A) )
  • P(A) is the prior, belief before the evidence.
  • P(B | A) is the likelihood of the evidence under A.
  • P(B) is the evidence or marginal likelihood, the normalising constant.
  • P(A | B) is the posterior.

The Developer block's fraud example is the expanded denominator, computed in counts. With prior 0.001, sensitivity 0.99 and false-positive rate 0.02, the posterior is 0.00099 / (0.00099 + 0.01998) = 0.0472.

Stating it in odds form makes the structure visible and removes the arithmetic:

posterior_odds  =  likelihood_ratio  ×  prior_odds
LR+             =  sensitivity / (1 - specificity)  =  0.99 / 0.02  =  49.5
prior_odds      =  1 / 999
posterior_odds  =  49.5 / 999  =  0.0496   ->  probability 0.0472

A likelihood ratio of 49.5 is a strong test. It is still not strong enough to overcome prior odds of 1 to 999. Evidence strength and base rate multiply; neither one alone decides the answer.

Independence is P(A and B) = P(A) P(B). Conditional independence given C is P(A, B | C) = P(A|C) P(B|C). The second does not imply the first, nor the reverse. Naive Bayes assumes conditional independence of features given the class. That assumption is false for text, and the classifier stays competitive anyway. The argmax of the posterior is robust to badly miscalibrated magnitudes (Domingos and Pazzani, 1997).

Distributions you will actually meet

DistributionSupportWhere it appears
Bernoulli(p){0,1}binary labels; sigmoid output head
Categorical(p_1..p_K){1..K}softmax output head; next-token distribution
Gaussian(mu, sigma**2)Rweight init, noise models, VAE latents, diffusion
Laplace(mu, b)Rthe noise model behind an L1 loss; DP mechanisms
Dirichlet(alpha)simplexconjugate prior over categoricals; topic models
Beta(a, b)[0,1]conjugate prior for Bernoulli; Thompson sampling

The Gaussian's dominance is not aesthetic. The central limit theorem states that for i.i.d. X_i with finite mean mu and variance sigma**2, the normalised sum converges in distribution:

sqrt(n) * ( mean(X) - mu ) / sigma   ->   Normal(0, 1)

Convergence is in distribution only, and the rate is O(1/sqrt(n)) by the Berry-Esseen bound. Heavy-tailed data with infinite variance is not covered at all, and gradient-norm distributions in deep networks are frequently heavy-tailed.

Losses are negative log-likelihoods

Maximum likelihood picks theta maximising sum_i log p(y_i | x_i; theta). Two standard losses drop straight out of two noise assumptions:

  • y | x ~ Normal(f(x; theta), sigma**2) with fixed sigma gives -log p = (y - f)**2 / (2 sigma**2) + const, so minimising it is mean squared error.
  • y | x ~ Categorical(softmax(f(x; theta))) gives -log p = -log p_correct, which is cross-entropy.

Cross-entropy decomposes as

H(p, q)  =  H(p)  +  KL(p || q)
KL(p || q)  =  sum_x p(x) log( p(x) / q(x) )   >=  0,  zero iff p == q

p is the true label distribution, q the model's. H(p) is fixed by the data, so minimising cross-entropy minimises the KL divergence from truth to model. KL is not symmetric and is not a metric. The direction you choose has real consequences. Forward KL is mass-covering, and produces blurry averages. Reverse KL is mode-seeking, and produces sharp but incomplete samples. That distinction drives the gap between VAE and adversarial objectives.

Calibration

A model is perfectly calibrated when P(Y = y | p_hat(y) = q) = q for all q. Measured with expected calibration error over M confidence bins:

ECE  =  sum_{m=1..M}  (|B_m| / n) * | acc(B_m) - conf(B_m) |

B_m is the set of predictions whose confidence falls in bin m. acc is its empirical accuracy, and conf its mean confidence.

Guo et al. (2017) documented that modern networks are systematically overconfident. The effect worsens with depth and width, even as accuracy improves.

Their fix is temperature scaling. Divide the logits by a single scalar T, fitted on a validation set by minimising negative log-likelihood. One parameter, accuracy unchanged since argmax is invariant to positive scaling, and ECE typically drops by an order of magnitude. It remains the best return-on-effort intervention available.

Caveats worth carrying: ECE is sensitive to binning scheme and is biased downward with few bins. Calibration under distribution shift degrades sharply (Ovadia et al., 2019). Calibration is also a marginal, population-level property. A model can be perfectly calibrated overall and badly miscalibrated on every subgroup.

Where probability enters modern systems

Sampling. Next-token selection draws from the categorical distribution after temperature, top-k or nucleus truncation (Holtzman et al., 2019). Truncation deliberately breaks calibration to improve perceived quality, which is a fair trade and worth naming as one.

Variational inference. The ELBO bounds an intractable log marginal likelihood:

log p(x)  >=  E_{q(z|x)}[ log p(x|z) ]  -  KL( q(z|x) || p(z) )

This is the VAE objective (Kingma and Welling, 2013), and the same decomposition underlies diffusion training objectives.

Uncertainty estimation. Deep ensembles (Lakshminarayanan et al., 2017) remain the strongest practical baseline. They beat MC dropout on calibration and on out-of-distribution detection.

Two kinds of uncertainty are worth separating. Aleatoric uncertainty is irreducible noise in the data. Epistemic uncertainty is reducible ignorance in the model. The split matters for active learning, and for deciding when a system should refuse to answer.

References

  • Kolmogorov, A. N. Grundbegriffe der Wahrscheinlichkeitsrechnung, 1933.
  • Jaynes, E. T. Probability Theory: The Logic of Science. Cambridge University Press, 2003.
  • Bishop, C. M. Pattern Recognition and Machine Learning. Springer, 2006. Chapters 1 to 2.
  • Murphy, K. P. Probabilistic Machine Learning: An Introduction. MIT Press, 2022.
  • Domingos, P., Pazzani, M. "On the Optimality of the Simple Bayesian Classifier under Zero-One Loss." Machine Learning 29, 103–130, 1997.
  • Guo, C., Pleiss, G., Sun, Y., Weinberger, K. Q. "On Calibration of Modern Neural Networks." ICML, 2017. arXiv:1706.04599
  • Lakshminarayanan, B., Pritzel, A., Blundell, C. "Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles." NeurIPS, 2017. arXiv:1612.01474
  • Ovadia, Y., et al. "Can You Trust Your Model's Uncertainty?" NeurIPS, 2019. arXiv:1906.02530
  • Holtzman, A., Buys, J., Du, L., Forbes, M., Choi, Y. "The Curious Case of Neural Text Degeneration." ICLR, 2020. arXiv:1904.09751
  • Kingma, D. P., Welling, M. "Auto-Encoding Variational Bayes." arXiv:1312.6114, 2013.

What to learn next