Semi-supervised learning
In one sentence Semi-supervised learning trains on a small labelled set plus a large unlabelled one, letting the cheap data amplify the expensive labels.
Updated
Semi-supervised learning combines a small amount of labelled data with a large amount of unlabelled data, using the unlabelled bulk to sharpen what the labels teach.
Think of learning to sort mangoes at a market. An expert spends one morning labelling a single crate for you: ripe, raw, spoilt. Then you work alone through a truckload. The expert's crate gave you the categories; the truckload teaches you the range — how much colours vary, where the borderline cases sit. The unlabelled bulk did not tell you any answers, yet you finish far better calibrated than the one crate alone could make you.
This matters because in most real projects, unlabelled data is nearly free while labels cost money and expert time — a hospital has a million scans and budget to have a radiologist mark two thousand.
The common mechanisms: self-training, where a model trained on the labelled set predicts labels for unlabelled data, and its confident predictions (pseudo-labels) are added to training; and consistency training, which pushes the model to give the same answer for an example and a lightly distorted copy of it. Both lean on the assumption that examples near each other deserve the same label. When that assumption fails, pseudo-label errors can compound — confident nonsense teaching more confident nonsense — so results are always checked against a properly labelled validation-set.
Where to go next
- Full lesson: Supervised learning
- Related terms: supervised-learning, unsupervised-learning, data-labelling, label