AI glossary

Clustering

In one sentence Clustering groups similar examples together without being told what the groups are — the algorithm discovers them.

By Updated

Clustering automatically groups similar examples together, without anyone defining the groups in advance.

Guests at a wedding sort themselves. Nobody assigns tables, yet an hour in, the college friends have found one corner, the office colleagues another, the aunties a third. Similarity did the organising. Clustering algorithms watch for the same effect in data: examples that sit near each other — by whatever features you measured — get grouped, and the groups emerge rather than being taught.

This is the flagship task of unsupervised-learning, because no labels exist. The most used algorithm is k-means; alternatives like DBSCAN find odd-shaped groups and flag points that belong nowhere — useful for outlier detection.

Where it earns money: customer segmentation ("these 40,000 shoppers behave alike — one campaign for them"), grouping support tickets by theme, organising news articles, and compressing image colours. On text, clustering embeddings groups documents by meaning rather than shared words.

The honest difficulties: you usually must choose the number of clusters, and the data rarely announces it. Results depend heavily on feature-scaling and on which features you measured in the first place. And the algorithm will return clusters even in structureless data — whether the clusters mean anything is a judgement the domain expert, not the algorithm, has to make.

Where to go next