Transfer learning
In one sentence Transfer learning reuses a model trained on one large task as the starting point for a related task with far less data.
Updated
Transfer learning takes a model trained on one big task and reuses its knowledge as the starting point for a different, related task.
A trained tabla player picking up the drum kit is not starting from zero. Rhythm, timing, hand independence — years of transferable skill carry over, and only the kit-specific technique needs learning. A complete beginner needs years; the tabla player needs months. Starting from accumulated skill beats starting from scratch.
In deep learning, the "accumulated skill" is concrete: the early layers of a network trained on millions of images learn edges, textures and shapes — vision itself, not cats specifically. Those layers transfer. The standard recipe: take a pretrained network, keep its layers, replace the final task-specific layer, and train briefly on your small dataset. Either freeze the borrowed layers and train the new head, or fine-tune everything gently with a low learning-rate.
The practical consequence is dramatic: tasks that would need millions of labelled examples become feasible with a few thousand. A crop-disease classifier can start from an ImageNet-trained backbone; a legal-document classifier starts from a pretrained language model.
This idea, scaled to its extreme, is the era we live in: pretraining one giant foundation-model once, then adapting it everywhere cheaply.
Where to go next
- Full lesson: Fine-tuning
- Related terms: fine-tuning, pretraining, foundation-model, distillation