Monitoring Models in Production
Concept drift
Concept drift is when the correct answer for the same input quietly changes, so a model can look perfectly calibrated on its inputs and still be wrong.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Concept drift is when the correct answer for the same situation quietly changes, even though the situation itself looks the same.
The analogy you have already lived
A parent moves abroad for ten years and eats mild food the whole time. They come home and taste their mother's cooking, cooked exactly the same way it always was.
It tastes far spicier to them than it used to. The chilli in the dish did not change. Their own sense of "spicy" did.
Someone using the parent's old spice judgment would now get every decision wrong. The dish itself never changed at all.
Why it exists
The previous lesson covered inputs changing shape. Concept drift is a different, sneakier problem.
The inputs can look completely normal. Same range, same average, same everything you would check with a drift score. What changed is the rule connecting an input to its correct answer.
A model has no way to notice this on its own. It keeps applying an old rule to inputs that look familiar, confident every time, and wrong more and more often.
How it works
BEFORE AFTER
ticket open > 5 min -> urgent ticket open > 5 min -> unusual, low priority
(team was slow, so a long wait (team got fast, so a long wait now means
really did mean "urgent") something odd stalled it, not urgency)
same input shape. opposite correct answer.Nothing about the ticket-length numbers changed. What changed is what a long wait means. That meaning is exactly the thing a drift score on the input cannot see.
A real example you have seen
Spam filters are the classic case. Spammers keep changing their wording to dodge whatever rule currently catches them.
An old, obvious scam pattern stops looking like one. A normal-sounding message starts carrying the scam instead. The words a filter watches for become the wrong words to watch.
The honest part
This one is genuinely hard to catch early, and that is normal, not a sign you are missing something obvious.
Data drift shows up in the inputs. Concept drift only shows up once you compare the model's predictions against what actually turned out to be true. Without those true answers, arriving late, concept drift can run for weeks in silence.
Remember this
- Concept drift is the input-to-answer rule changing, not the inputs themselves.
- Input checks like PSI can stay completely calm while concept drift does real damage.
- It is only visible once you compare predictions against real outcomes, which usually arrive late.
What to learn next
- Working with delayed labels — why concept drift is so often caught late, and how to shrink that delay.
- Monitoring by segment — catching a drift that only hit one slice of traffic.
- Monitoring and model drift — the wider picture both drift lessons zoom into.
Developer — Code and libraries.
Setup
pip install numpy scikit-learnBuilding a case where drift checks stay quiet
One feature. Its distribution never changes between the two periods. Only the rule connecting it to the right answer flips.
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
rng = np.random.RandomState(0)
def psi(expected, actual, buckets=10):
breakpoints = np.quantile(expected, np.linspace(0, 1, buckets + 1))
breakpoints[0], breakpoints[-1] = -np.inf, np.inf
e_counts = np.histogram(expected, bins=breakpoints)[0] / len(expected)
a_counts = np.histogram(actual, bins=breakpoints)[0] / len(actual)
e_counts = np.clip(e_counts, 1e-4, None)
a_counts = np.clip(a_counts, 1e-4, None)
return float(np.sum((a_counts - e_counts) * np.log(a_counts / e_counts)))
# One feature: minutes since a support ticket was opened.
# Its distribution never changes -- same range, before and after.
x_before = rng.uniform(0, 10, 1000).reshape(-1, 1)
x_after = rng.uniform(0, 10, 1000).reshape(-1, 1)
# BEFORE: tickets open longer than 5 minutes were the urgent ones.
y_before = (x_before[:, 0] > 5).astype(int)
# AFTER: the support team got faster. Now anything still open past
# 5 minutes is the unusual, low-priority case that stalled -- the
# meaning of "long open time" flipped. Same input range, opposite rule.
y_after = (x_after[:, 0] < 5).astype(int)
model = LogisticRegression().fit(x_before, y_before)
print(f"PSI on the input feature: {psi(x_before[:, 0], x_after[:, 0]):.3f} (no data drift)")
print(f"Accuracy on pre-drift holdout: {accuracy_score(y_before, model.predict(x_before)):.3f}")
print(f"Accuracy on post-drift data: {accuracy_score(y_after, model.predict(x_after)):.3f}")PSI on the input feature: 0.025 (no data drift) Accuracy on pre-drift holdout: 1.000 Accuracy on post-drift data: 0.000
This is exact output from this seeded script, not an estimate. The rule flip here is deliberately total, to make the effect visible in twenty lines. Real concept drift is almost always partial, and the accuracy drop is gentler and slower.
Line-by-line walkthrough
x_before and x_after are drawn from the identical rng.uniform(0, 10, ...) call shape. A drift score comparing them will correctly report "nothing changed", because nothing did.
y_before and y_after use opposite comparisons on the same feature. That single flipped operator (> becomes <) is the entire concept drift in this example.
The model trained on x_before, y_before reaches perfect accuracy on its own kind of data, and zero accuracy on x_after, y_after. It is not confused. It is confidently, completely wrong.
Common mistakes
Trusting a clean PSI score as proof the model is fine. PSI only checks inputs. It has no way to see a broken input-to-answer rule, by design.
Waiting for accuracy to be the first signal you check. By the time labelled accuracy is available, real decisions were already made on the old rule. See working with delayed labels for how to shorten that gap.
Assuming a drift is total, like the example above. Most real concept drift affects one segment of traffic, or shifts gradually over months. Monitoring by segment catches the partial cases that a single overall accuracy number hides.
Try it yourself
Change y_after to flip only for x_after[:, 0] > 8, leaving the rest of the rule unchanged. Rerun, and watch accuracy fall by far less than to zero. That partial flip is closer to what real concept drift looks like.
What to learn next
- Working with delayed labels — why concept drift is so often caught late, and how to shrink that delay.
- Monitoring by segment — catching a drift that only hit one slice of traffic.
- Monitoring and model drift — the wider picture both drift lessons zoom into.
Researcher — Mathematics and papers.
A taxonomy of drift patterns
Gama et al. (2014) formalise four temporal shapes for concept drift, distinguished by how $P(y \mid x)$ changes over time:
- Sudden — an abrupt, step-function change at one point in time.
- Gradual — a period where both the old and new concept generate observations, with the new one's share rising.
- Incremental — many small steps between the old and new concept, with no single point of full transition.
- Recurring — a concept that returns after some interval, as with seasonal fraud patterns.
The example in the developer tab is the sudden case, chosen for legibility. Detection difficulty rises sharply from sudden to incremental, because incremental drift can be statistically indistinguishable from noise over any short window.
Detection without waiting for full retraining
Page-Hinkley test — a sequential test on a running error statistic, raising an alarm when cumulative deviation from the running mean exceeds a threshold, with a tunable sensitivity-versus-false-alarm trade-off.
ADWIN (Adaptive Windowing), Bifet and Gavaldà (2007), maintains a variable-length window of recent error values and splits it when two sub-windows differ in mean by more than a bound derived from the Hoeffding inequality. The window's own length becomes an implicit drift signal.
DDM / EDDM monitor the online error rate and its standard deviation directly, flagging both a "warning" and a "drift" level, so retraining can be triggered before performance fully collapses.
Distinguishing concept drift from data drift formally
For a model trained under $P_{\text{train}}(x, y) = P_{\text{train}}(y \mid x) \, P_{\text{train}}(x)$, data drift is a change in $P(x)$, and concept drift is a change in $P(y \mid x)$. Both can occur independently or together, and only the joint distribution is directly observable from raw logs — $P(y \mid x)$ requires labels, which is why concept drift detection is fundamentally label-hungry in a way data drift detection is not.
Complexity and cost
ADWIN's amortised update cost is $O(\log W)$ per new observation for a window of size $W$, using an exponential histogram internally rather than storing the full window. Page-Hinkley is $O(1)$ per observation. Both are designed for streaming use, not batch recomputation.
Papers
- Gama, Žliobaitė, Bifet, Pechenizkiy and Bouchachia, A Survey on Concept Drift Adaptation, ACM Computing Surveys 2014 — dl.acm.org/doi/10.1145/2523813
- Bifet and Gavaldà, Learning from Time-Changing Data with Adaptive Windowing, SDM 2007
- Gonçalves et al., A Comparative Study on Concept Drift Detectors, Expert Systems with Applications 2014
What to learn next
- Working with delayed labels — why concept drift is so often caught late, and how to shrink that delay.
- Monitoring by segment — catching a drift that only hit one slice of traffic.
- Monitoring and model drift — the wider picture both drift lessons zoom into.