Linear discriminant analysis
LDA finds the viewing angle that pulls class averages far apart while keeping each class tight — one move that classifies and compresses at the same time.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Linear discriminant analysis finds the viewing angle from which different groups look most separated.
Think of a group photo where two families stand mixed together on some steps. From the front, everyone overlaps. But walk to the side, crouch a little, and from one particular angle the two families fall into two clean clusters. Same people, same steps — the separation came from choosing where to stand.
Linear discriminant analysis (LDA) computes that viewpoint from data. It then classifies new points by which cluster they land in, seen from there.
Why it exists
Data often has many measurements per example — length, width, weight, colour, and so on. Most of those measurements mix the groups together. Some hide a clean separation, but only when combined in the right proportions.
LDA was invented to find the combination worth looking at. It asks for two things at once: push the group averages far apart, and keep each group tightly bunched around its own average. Distance between groups is worthless if each group sprawls; tightness is worthless if the averages sit on top of each other. LDA maximises the ratio of the two.
The bonus: the viewing angle is useful on its own. Squeezing many measurements down to a few tell-tale directions makes data easier to plot, store and feed to other models.
How it works
front view (bad angle): side view (LDA's angle):
● ○ ● ○ ○ ● ●●● ○○○
○ ● ○ ● ● ○ → ●●● ○○○
groups overlap far apart, each tightA real example you have seen
Voice assistants distinguishing speakers, and early face-recognition systems ("Fisherfaces"), used LDA to find the few directions in which people differ most. Banks separate "will repay" from "will default" on a handful of financial ratios. They have used discriminant analysis for this since the 1960s. That is decades before anyone said "machine learning".
Remember this
- LDA finds the angle of view that separates groups best.
- Good separation = averages far apart AND groups tightly bunched — the ratio matters.
- One tool, two jobs: a classifier and a dimensionality reducer.
What to learn next
- Statistics — means, spreads and the language LDA is written in.
- Naive Bayes — a different generative classifier, built on counting.
- Logistic regression — the discriminative model sharing LDA's linear boundary.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7.2 on CPU.
Three flowers, four measurements, two directions
from sklearn.datasets import load_iris
from sklearn.discriminant_analysis import LinearDiscriminantAnalysis
from sklearn.model_selection import cross_val_score
X, y = load_iris(return_X_y=True) # ships inside scikit-learn, no download
lda = LinearDiscriminantAnalysis()
acc = cross_val_score(lda, X, y, cv=5).mean()
print(f"5-fold accuracy: {acc:.3f}")
proj = lda.fit(X, y).transform(X) # 4 measurements squeezed to 2 axes
print("projected shape:", proj.shape)
print("class separation carried by each axis:", lda.explained_variance_ratio_.round(3))
for cls in (0, 1, 2):
print(f"class {cls} mean on axis 1: {proj[y == cls, 0].mean():6.2f}")5-fold accuracy: 0.980 projected shape: (150, 2) class separation carried by each axis: [0.991 0.009] class 0 mean on axis 1: 7.61 class 1 mean on axis 1: -1.83 class 2 mean on axis 1: -5.78
The walkthrough
98% accuracy from a model with no knobs. No learning rate, no tree depth, no C. LDA fits class averages and spreads in closed form — one shot, no iterations. As a fast, honest baseline before anything fancier, it is hard to beat.
Why 2 axes from 4 measurements? With 3 classes, all the between-average geometry lives in a flat sheet through 3 points. LDA can produce at most classes - 1 directions — here 2, whatever the number of input features. This cap is structural, not a setting.
explained_variance_ratio_ reads [0.991, 0.009]. The first direction carries 99% of the separation; the second is almost decorative. For plotting, one axis nearly suffices — and the class means on it confirm the story: 7.61, -1.83, -5.78. Setosa sits far from the other two, versicolor and virginica are neighbours. Anyone who has plotted iris knows that picture; LDA found it from the ratio alone.
transform vs predict. transform gives coordinates in the discriminant directions — LDA as compressor, a supervised cousin of PCA. predict classifies. Same fitted object, two products.
Common mistakes
Using PCA when labels exist and separation is the goal. PCA finds directions of biggest overall spread, labels unseen; the most spread-out direction can be exactly the one where classes overlap. If you have labels and want separation, LDA is the tool built for the question.
Expecting more than classes - 1 components. LinearDiscriminantAnalysis(n_components=3) with 3 classes raises ValueError. The ceiling is baked into the geometry.
Ignoring the shared-spread assumption. LDA assumes every class scatters in roughly the same shape and size. When one class is much more spread out, its boundary gets misplaced — QuadraticDiscriminantAnalysis drops the assumption at the cost of many more parameters.
Many features, few samples. With more features than samples, the within-class spread estimate becomes singular and everything gets shaky. Use LinearDiscriminantAnalysis(solver="lsqr", shrinkage="auto") — regularised spread estimation designed for exactly this regime.
Try it yourself
Run PCA(n_components=2) on the same iris data and print each class's mean on PCA's first axis. Compare with LDA's means above — which axis separates versicolor from virginica better? Then try shrinkage="auto" with solver="lsqr" and check whether accuracy moves.
What to learn next
- Statistics — means, spreads and the language LDA is written in.
- Naive Bayes — a different generative classifier, built on counting.
- Logistic regression — the discriminative model sharing LDA's linear boundary.
Researcher — Mathematics and papers.
Fisher's criterion
For classes $c = 1, \dots, K$ with means $\mu_c$, overall mean $\mu$, and $n_c$ samples each, define the scatter matrices:
$$ S_W = \sum_{c=1}^{K} \sum_{i \in c} (x_i - \mu_c)(x_i - \mu_c)^\top, \qquad S_B = \sum_{c=1}^{K} n_c (\mu_c - \mu)(\mu_c - \mu)^\top $$
Where:
- $S_W$ — within-class scatter: how much each class sprawls around its own mean.
- $S_B$ — between-class scatter: how far class means sit from the global mean.
- $x_i \in \mathbb{R}^d$ — a sample; $d$ — number of features.
Fisher's criterion seeks the direction $w$ maximising:
$$ J(w) = \frac{w^\top S_B \, w}{w^\top S_W \, w} $$
The maximisers are the top eigenvectors of $S_W^{-1} S_B$. Since $\operatorname{rank}(S_B) \leq K - 1$, at most $K - 1$ informative directions exist — the cap met in the code. For $K = 2$ the solution is closed-form: $w \propto S_W^{-1}(\mu_1 - \mu_0)$ (Fisher, 1936, The use of multiple measurements in taxonomic problems — the paper that also gave the world the iris dataset).
The generative view
Model each class as a Gaussian with its own mean but a shared covariance $\Sigma$: $x \mid c \sim \mathcal{N}(\mu_c, \Sigma)$. The log-posterior ratio between two classes is then linear in $x$, giving discriminant functions:
$$ \delta_c(x) = x^\top \Sigma^{-1} \mu_c - \tfrac{1}{2} \mu_c^\top \Sigma^{-1} \mu_c + \log \pi_c $$
with $\pi_c$ the class prior. Classify by the largest $\delta_c$. Dropping the shared-$\Sigma$ assumption yields quadratic boundaries (QDA), at the price of $K$ separate covariance estimates. LDA's boundary equals logistic regression's form, but LDA estimates it generatively — more efficient when the Gaussian assumption holds, brittler when it fails.
Complexity and regularisation
Fitting costs $O(nd^2 + d^3)$ — scatter accumulation plus an eigendecomposition; prediction is $O(Kd)$. When $d \gtrsim n$, $S_W$ is singular. Shrinkage LDA replaces $\hat\Sigma$ with $(1-\lambda)\hat\Sigma + \lambda \bar\sigma^2 I$, with $\lambda$ set analytically by Ledoit and Wolf (2004) — this is shrinkage="auto".
Descendants
- Fisherfaces (Belhumeur et al., 1997): PCA then LDA for face recognition, dominant pre-deep-learning.
- Kernelised: Kernel Fisher Discriminant (Mika et al., 1999) via the kernel trick.
- The "maximise between over within" objective survives inside modern metric learning and speaker-embedding losses; PLDA (probabilistic LDA, Ioffe, 2006) remains standard scoring in speaker verification.
What to learn next
- Statistics — means, spreads and the language LDA is written in.
- Naive Bayes — a different generative classifier, built on counting.
- Logistic regression — the discriminative model sharing LDA's linear boundary.