AI glossary

Validation set

In one sentence The validation set is the data you tune your model against during development, kept separate from both training data and the final test set.

By Updated

A validation set is held-out data used during development to compare settings and decide when to stop — so the final test set stays untouched.

Why three sets and not two? Think of a cook preparing for a competition. The training set is practice in the kitchen. The validation set is the family: they taste every trial dish, and the cook adjusts to their feedback. The test set is the judges, who taste exactly once. If the family and the judges were the same people, the cook would slowly tune the dish to those specific palates — and the competition score would mean nothing.

That slow tuning is the danger in numbers. Every time you pick a hyperparameter because it scored best on some dataset, you fit yourself to that dataset a little. Doing this against the test set quietly inflates the final score. The validation set exists to absorb this damage, leaving the test set as an honest, once-only measurement.

Day to day, the validation set is the busiest of the three. It drives early-stopping, chooses between model variants, and draws the second line on every training curve you will ever read. When data is too scarce to spare a fixed slice, cross-validation plays the same role by rotation.

Where to go next