Dimensionality Reduction

Principal component analysis

PCA finds the few directions along which your data varies most and rewrites every row using only those, compressing hundreds of columns with minimal loss.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

PCA compresses your columns by finding the few directions in which the data actually varies, and dropping the rest.

Hold a wireframe model of a bird up against the evening sun and look at its shadow on the wall. From most angles the shadow is a blob. But there is one angle where the shadow shows the wings, the beak, the whole story. One flat picture keeps almost everything worth knowing about a solid object.

PCA is the machine that finds that angle. Your data is the solid object, living in as many directions as you have columns. PCA rotates it, looking for the viewpoint from which the data's shape spreads out the most, and photographs it from there.

Why it exists

The curse of dimensionality showed that extra columns thin out your data and break distance. But columns are rarely independent stories. Height and weight move together. Income, spending and savings move together. A hundred columns often contain perhaps five genuinely separate patterns, each echoed across many columns.

PCA hunts those underlying patterns. Each one it finds is called a principal component — a direction through the data, built as a weighted blend of the original columns. The first component is the single direction that captures the most variation. The second captures the most of what remains, at right angles to the first. And so on, each new direction mopping up leftovers.

How it works

150 flowers, 4 measurements each:

petal length ─┐
petal width  ─┤   PCA finds: "these four move together
sepal length ─┤   in ONE main pattern (overall size),
sepal width  ─┘   plus one smaller pattern (shape)"

4 columns  ──────>  2 columns, keeping ~96% of the variation

You choose how many components to keep. Keep two and every row is rewritten as two numbers: its position along pattern one and pattern two. The compression is honest — PCA tells you exactly what fraction of the variation survived.

A real example you have seen

Every personality quiz does this by hand. Hundreds of yes/no questions get boiled down to five scores: openness, extraversion and so on. This works because the hundreds of answers move in a handful of underlying patterns. PCA is the automatic version: give it the raw answers, and directions very like the "big five" fall out.

Remember this

  • A principal component is a direction of maximum variation — a blend of your columns.
  • Components come ranked, each one explaining a stated share of the variation.
  • PCA sees straight-line patterns; curved ones need kernel PCA or friends.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7.2.

Four measurements, two that matter

The iris dataset ships inside sklearn — 150 flowers, 4 measurements, no download.

pca.py
from sklearn.datasets import load_iris
from sklearn.decomposition import PCA
from sklearn.preprocessing import StandardScaler

X, y = load_iris(return_X_y=True)            # 150 flowers, 4 measurements
X = StandardScaler().fit_transform(X)        # PCA needs comparable scales

pca = PCA(n_components=2).fit(X)
X2 = pca.transform(X)

print("shape before:", X.shape, " after:", X2.shape)
print("variance kept per direction:", pca.explained_variance_ratio_.round(3))
print("total kept:", pca.explained_variance_ratio_.sum().round(3))
print("direction 1 recipe:", pca.components_[0].round(2))
print("first flower, compressed:", X2[0].round(2))
Output
shape before: (150, 4)  after: (150, 2)
variance kept per direction: [0.73  0.229]
total kept: 0.958
direction 1 recipe: [ 0.52 -0.27  0.58  0.56]
first flower, compressed: [-2.26  0.48]

The walkthrough

explained_variance_ratio_ is the honesty meter. Component one alone carries 73% of all variation; two together carry 95.8%. Four columns became two at a cost of 4.2% of the variation. That number is your receipt — always print it.

The recipe line is readable. Component one weights sepal length (0.52), petal length (0.58) and petal width (0.56) positively and sepal width negatively (-0.27). In words: "big-petalled, long-sepalled flowers versus small ones" — an overall size-and-species axis. Components are blends, and the weights say what each blend measures.

Scale first, always. Without standardising, the column with the biggest units wins the "most variation" contest by cheating. PCA on raw rupees-and-years data returns a first component that is essentially "the rupees column". Scale, then fit.

Round trips measure what you lost. pca.inverse_transform(X2) maps the 2 numbers back to 4. The difference from the original is exactly the discarded 4.2% — useful for compression demos and anomaly detection: rows that reconstruct badly do not follow the main patterns.

Common mistakes

Fitting PCA on all rows before the train-test split. The components are learned from data, so learning them from test rows leaks. Fit on train, transform both — the same discipline as every fitted step.

Choosing components by ritual. "Keep 95%" is folklore, not law. For visualisation you keep 2. For feeding a classifier, treat n_components as a hyperparameter and cross-validate. PCA(n_components=0.95) — passing a float — picks the count reaching that share for you.

Reading components as physical causes. Component one is the direction of most variance — a mathematical summary, not a mechanism. The size axis in iris does not prove a "size gene". Naming components is interpretation, and over-naming them is the field's oldest sin.

Running PCA on sparse one-hot matrices. PCA centres the data first, which destroys sparsity and can explode memory on wide matrices. That situation belongs to truncated SVD, which skips the centring.

Expecting PCA to help classification automatically. PCA is unsupervised — it never sees the labels. The directions of most variance are sometimes not the directions that separate classes. If accuracy drops after PCA, that is why.

Try it yourself

Set n_components=4 and print all four variance ratios. Then reconstruct with 1, 2 and 3 components and print the worst per-flower reconstruction error for each — watch how quickly "good enough" arrives.

What to learn next

Researcher — Mathematics and papers.

Two equivalent derivations

Centre the data matrix X (n rows, d columns). PCA seeks orthonormal directions w maximising projected variance:

w_1 = argmax over ||w|| = 1 of w^T S w, S = (1/n) X^T X

Where S is the sample covariance matrix. The maximiser is the leading eigenvector of S; component k is the k-th eigenvector, with eigenvalue lambda_k equal to the variance captured. explained_variance_ratio_ is lambda_k / sum of all lambdas.

Equivalently, PCA minimises reconstruction error: among all rank-k linear projections, the top-k eigenbasis minimises the sum of squared distances from points to their projections. Maximising kept variance and minimising lost variance are the same optimisation read from opposite ends. Pearson (1901) derived it the second way; Hotelling (1933) the first.

SVD is how it is actually computed

Numerically, sklearn never forms S. It takes the singular value decomposition X = U Sigma V^T, where V's columns are the principal directions and lambda_k = sigma_k^2 / n. SVD on X is better conditioned than eigendecomposition of X^T X (condition number squared in the latter). Full SVD costs O(min(n d^2, n^2 d)); for large problems sklearn switches to randomized SVD (Halko, Martinsson and Tropp, 2011) — random projection to a small subspace, then exact SVD there — giving near-optimal rank-k approximations in O(n d k) with high probability.

Choosing k with more than folklore

  • Scree/elbow inspection and cumulative variance thresholds are heuristics with no inferential content.
  • Under a spiked covariance model, random-matrix theory gives a principled cut: eigenvalues below the Marchenko-Pastur bulk edge (1 + sqrt(d/n))^2 (for unit noise) are indistinguishable from noise — components beyond that are not signal.
  • Probabilistic PCA (Tipping and Bishop, 1999) recasts PCA as maximum likelihood in a Gaussian latent model x = W z + mu + eps, enabling likelihood-based model selection, EM fitting with missing data, and Bayesian treatment of k.

Boundaries of the method

PCA is optimal only in the L2, linear, Gaussian-flavoured world. Heavy-tailed noise breaks it (robust PCA, Candes et al., 2011, decomposes into low-rank plus sparse). Nonlinear structure hides from it (kernel PCA; autoencoders are its learned nonlinear descendants). Independence, not decorrelation, needs ICA. And its components are dense blends of all features; sparse PCA trades captured variance for interpretable loadings.

What to learn next