Distillation
In one sentence Distillation trains a small student model to imitate a large teacher model, keeping much of the capability at a fraction of the cost.
Updated
Distillation trains a small "student" model to reproduce the outputs of a large "teacher" model, compressing capability into a cheaper package.
A master chef cannot be photocopied, but she can train an apprentice — and crucially, the apprentice learns from her finished judgements, not from the raw cookbooks she spent decades studying. Watching the master season a dish teaches more per hour than re-reading every book she ever read. The student model gets the same shortcut: it learns from the teacher's answers instead of from raw data alone.
The classical trick (Hinton et al., 2015) uses the teacher's soft outputs. A hard label says "this is a 7". The teacher's full probability spread says "0.90 seven, 0.08 one, 0.01 nine" — encoding that sevens resemble ones. Those between-class relationships, invisible in hard labels, are exactly the knowledge worth transferring, and the student trains to match them.
In the LLM era, distillation increasingly means generating synthetic-data: the teacher writes instruction-and-answer pairs, and the student fine-tunes on them. Many strong small open models are trained substantially this way from frontier-model outputs — and most commercial API terms restrict using outputs to train competitors, a legal edge worth knowing.
Distillation compresses by re-training; quantization compresses the same weights into fewer bits. They stack: a distilled-then-quantized model is the standard recipe for phones and edge devices.
Where to go next
- Full lesson: Fine-tuning
- Related terms: quantization, transfer-learning, synthetic-data, inference