Unsupervised learning
In one sentence Unsupervised learning finds structure in data that has no answers attached — groups, patterns, and oddities.
Updated
Unsupervised learning is finding structure in data that comes with no correct answers — no labels, only the data itself.
Hand someone a sack of mixed buttons and say "organise these". Nobody tells them the categories. They will still produce neat piles — by colour, by size, by number of holes — because the structure is in the buttons. Unsupervised learning is the algorithmic version: given only inputs, find the groups, the patterns, and the pieces that fit nowhere.
The main jobs:
clustering group similar things customer segments, topic groups
dimensionality reduction compress and visualise PCA to plot 100 columns in 2D
anomaly detection flag what fits no pattern fraud, failing machinesThe catch is that "similar" is genuinely ambiguous. The button-sorter could be organising by colour when you wanted size, and neither is wrong. With no label to check against, evaluating unsupervised results is partly judgement — do the customer segments mean something to the marketing team? — which makes this family harder to use well than supervised-learning, not easier.
A note on the modern landscape: LLM pretraining is often loosely called unsupervised because no human writes labels. The precise term is self-supervised-learning — the answers exist, manufactured from the data itself.
Where to go next
- Full lesson: Unsupervised learning
- Related terms: clustering, pca, supervised-learning, outlier