AI glossary

Train/test split

In one sentence A train/test split holds back part of your data so the model is scored on examples it has never seen.

By Updated

A train/test split divides your data into a part the model learns from and a part kept hidden until scoring time.

A teacher who sets the exam using the exact questions from the homework learns nothing about her students — only whether they can memorise. To test understanding, the exam must contain fresh questions on the same material. The test set is that fresh exam: data from the same source, never shown during training.

The mechanics are one line, with an 80/20 split as the common default:

python
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

The entire value of the test set is its innocence, and losing that innocence is the classic beginner failure. Fitting a scaler on all the data before splitting leaks test statistics into training. Tuning hyperparameters against the test score, run after run, slowly turns the exam into homework. Both are forms of data-leakage, and both produce scores that collapse in production.

The disciplined setup uses three parts: train to fit, a validation-set to tune against, and a test set opened once, at the end. For time-ordered data, always split by time — train on the past, test on the future — because a random split lets the model peek ahead.

Where to go next