AI glossary

PCA (principal component analysis)

In one sentence PCA compresses many correlated features into a few new axes that keep most of the variation in the data.

By Updated

Principal component analysis, or PCA, finds the directions along which your data varies most, and re-describes each example using only those few directions.

Photographing a cricket bat is the intuition. The bat is a 3D object, but one photo from the right angle — side on — captures almost everything: length, blade shape, handle. A photo taken end-on captures almost nothing. PCA finds the "side on" angles for your data: the viewpoints that preserve the most variation, so you can drop the uninformative ones.

Mechanically, PCA finds new axes called principal components. The first is the single direction of greatest spread in the data. The second is the direction of greatest remaining spread at right angles to the first, and so on. Data with 100 correlated columns can often be summarised by 10 components with little loss — the components soak up the redundancy.

What it is used for: plotting high-dimensional data in 2D to eyeball clustering structure, speeding up downstream models, and de-noising, since low-variance directions are often noise. Two rules of use: apply feature-scaling first, or big-unit columns hijack the components; and remember components are linear blends of features, so sharply curved structure escapes it — that is where autoencoders and UMAP come in.

Where to go next