KNN (k-nearest neighbours)
In one sentence KNN classifies a new example by finding the k most similar known examples and letting them vote.
Updated
K-nearest neighbours, or KNN, predicts by looking up the k most similar past examples and taking their majority vote (or their average, for numbers).
It is how you judge an unfamiliar restaurant: find the three most similar places you already know — same cuisine, same price band, same neighbourhood — and assume this one is about as good. No theory of restaurants required. Your memory of past cases is the model.
KNN is unusual because there is no training step. The model memorises the entire training set, and all work happens at prediction time: measure the distance from the new point to every stored point, take the k closest, let them vote.
new patient → find 5 most similar past patients
4 recovered quickly, 1 did not
→ predict: recovers quicklyIts strengths are honesty and simplicity — it can fit any decision boundary shape, and its answers come with visible precedents ("here are the five cases behind this call"). Its weaknesses are the mirror image. Prediction is slow at scale, since every query scans the data. It is very sensitive to feature-scaling — unscaled columns silently dominate the distance. And in very high dimensions, "nearest" loses meaning as all points drift equally far apart.
The idea has a huge modern afterlife: vector-database search over embeddings is nearest-neighbour lookup, made fast with approximate indexes.
Where to go next
- Full lesson: Classification
- Related terms: classification, feature-scaling, vector-database, embedding