Training data
In one sentence Training data is the collection of examples a model learns from — its quality and coverage set the ceiling on everything the model can do.
Updated
Training data is the set of examples a model learns its behaviour from. Whatever is in it, the model absorbs — patterns, gaps, and mistakes alike.
A child raised in a kitchen learns that household's cooking: the recipes, the spice levels, and also the bad habits. The child cannot learn dishes the household never made. Models eat the same way. A speech model trained on news anchors struggles with rural accents it never heard. The model is, to a first approximation, a compressed reflection of its diet.
This is why practitioners repeat "garbage in, garbage out" like a prayer. Model architecture gets the glamour, but in applied work the highest-leverage questions are about the data: Is it representative — does it cover the accents, lighting conditions and customer types the model will actually meet? Are the labels right? Are there duplicates quietly inflating one pattern? Is anything in it something you are not licensed or ethical to use? Does it contain the answer by accident (data-leakage)?
Scale matters too, on a rough ladder: classical models on tables want thousands of rows; training image networks from scratch wants millions; LLM pretraining consumes trillions of tokens. Transfer-learning exists precisely to let you skip to the front of that ladder.
And the job never ends: the world drifts away from any frozen snapshot — see data-drift — so production training sets are living things, refreshed and re-audited.
Where to go next
- Full lesson: What is machine learning?
- Related terms: label, data-labelling, data-drift, train-test-split