AI glossary

Data drift

In one sentence Data drift is when the real-world data a deployed model receives slowly stops resembling the data it was trained on.

By Updated

Data drift is the gradual change of live input data away from what the model saw during training, quietly eroding its accuracy.

A weather-wise farmer's instincts are trained on decades of local patterns. Shift the monsoon by three weeks, year after year, and those instincts mislead — not because the farmer forgot anything, but because the world moved. Deployed models age the same way. The model is frozen; reality is not.

Drift arrives from everywhere: customer behaviour shifts after a festival season, a phone camera upgrade changes image statistics, a new competitor changes who visits your site, a form redesign changes how a field is filled. A fraud model trained pre-UPI meets a payments world it has never seen.

Two flavours deserve distinct names. Data drift (or covariate shift): the inputs change — new vocabulary, new demographics. Concept drift: the relationship between input and answer changes — the same transaction pattern that was safe in 2023 is fraudulent in 2026. The second is nastier, because inputs can look perfectly normal while the meaning underneath has moved.

The defence is monitoring, the heart of production MLOps. Compare live input distributions against training statistics, track prediction confidence over time, and — where labels eventually arrive — track live accuracy. When drift crosses a threshold, retrain on fresh data. The unglamorous truth: models in production are not finished artefacts but perishable goods.

Where to go next