MLOps

Monitoring and model drift

A model gets worse over time because the world changes around it. Monitoring is how you notice, and it has to work before the true answers arrive.

Read these first

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. Why this is not the same as "the code broke"
  4. The three ways it goes wrong
  5. The problem that makes all of this hard
  6. What you can watch today
  7. A real example you have lived through
  8. The honest part
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A model gets worse over time because the world it learned about keeps changing. Monitoring is how you find out before your users do.

The analogy you have already lived

You had a shortcut. A back lane that got you home ten minutes faster, and you used it for two years.

Then a new building opened on that road. Now the lane is worse than the main road. Nothing about your route changed. The world around it did.

Your model is that shortcut. It learned the world as it was on the day you trained it. The world did not agree to stay still.

Why this is not the same as "the code broke"

A broken service is loud. It returns errors, alarms go off, somebody gets a phone call at midnight.

A drifted model is silent. It answers every request. Every answer looks reasonable. The service is green on every dashboard.

It is only wrong.

That is why monitoring for machine learning needs its own lesson. Think about what you would normally watch: is it up, is it fast, is it erroring. All three stay perfectly green while the model quietly stops working.

The three ways it goes wrong

The people changed. Your loan model learned from customers earning around forty thousand. A campaign brings in younger customers earning less. The model never saw them, so its answers about them are guesses. This is called data drift.

The rules changed. Fraud used to look like large late-night transfers. Fraudsters read that too, and now send many small daytime ones. The same input now means something different. This is called concept drift, and it is the harder one.

Your data pipeline changed. Somebody upstream renamed a column, or started sending amounts in rupees where they used to send thousands. This is not really drift at all. It is a bug, it is very common, and drift monitoring catches it first.

The problem that makes all of this hard

You want to watch accuracy. You usually cannot.

To measure accuracy you need the true answer. For a loan, the true answer arrives when the customer repays or does not — in eighteen months. For a fraud alert, when the customer disputes it, or never. For a delivery time estimate, in an hour.

So you would find out your model broke in January, in the following July.

The way out is to watch things you can see today.

What you can watch today

                           can you see it today?
  the numbers going IN            yes  -->  are they the same shape as before?
  the answers coming OUT          yes  -->  is the mix of answers changing?
  how fast, how many errors       yes  -->  normal software monitoring
  whether it was RIGHT            no   -->  arrives weeks or months later

The middle two are the ones people forget, and they are free.

Suppose your model approved six out of ten applications all year. This week it approves four out of ten. Something has changed.

You do not know yet whether the model is wrong. You know to look today, instead of in July.

A real example you have lived through

Spam filters face the sharpest version of this. The people on the other side are deliberately changing their behaviour to defeat your model, every week.

That is why your spam filter is retrained constantly, and why a filter left alone for a year becomes useless. It is drift with an opponent.

Something similar happened at a huge scale in 2020. Almost every model predicting shopping, travel or footfall broke in the same month. Human behaviour had changed faster than any training set could describe.

The honest part

Drift alerts are noisy, and this part is genuinely difficult.

With enough data, a statistical test will report a "significant" change in something almost every week. Most of those changes do not affect your model at all. Teams that alert on every one of them stop reading alerts within a month, which is worse than having none.

Drift is a signal to investigate, not a signal to retrain. The judgement about which changes matter is real work, and no tool does it for you.

Remember this

  • Models decay because the world changes, not because the code broke.
  • The failure is silent — every ordinary dashboard stays green.
  • Watch the inputs and the answers today, because the truth arrives far too late.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scipy

Both come with scikit-learn already, so you very likely have them. Nothing here downloads anything.

Two measures, one control feature

The example compares a reference sample (what the model was trained on) against a live sample (what is arriving now). The age feature is included deliberately as a control. It did not move, so a good detector must stay quiet about it.

drift.py
import numpy as np
from scipy.stats import ks_2samp

rng = np.random.RandomState(0)

# What the model saw during training: monthly income in thousands.
reference = rng.normal(40, 12, 5000)

# Three months later. Marketing ran a campaign aimed at younger customers,
# so the people arriving now earn less. Nobody told the ML team.
live = rng.normal(33, 12, 5000)

# A feature that did not move, as a control.
ref_age = rng.normal(38, 9, 5000)
live_age = rng.normal(38, 9, 5000)


def psi(expected, actual, bins=10):
    """Population Stability Index. Bin edges come from the reference data,
    because the reference is what the model actually learned."""
    edges = np.quantile(expected, np.linspace(0, 1, bins + 1))
    edges[0], edges[-1] = -np.inf, np.inf
    e = np.histogram(expected, edges)[0] / len(expected)
    a = np.histogram(actual, edges)[0] / len(actual)
    e = np.clip(e, 1e-6, None)          # a zero bin would make the log blow up
    a = np.clip(a, 1e-6, None)
    return float(np.sum((a - e) * np.log(a / e)))


for name, ref, now in [("income", reference, live), ("age", ref_age, live_age)]:
    ks = ks_2samp(ref, now)
    print(f"{name:<7} mean {ref.mean():5.1f} -> {now.mean():5.1f}   "
          f"PSI {psi(ref, now):.3f}   KS {ks.statistic:.3f}   p {ks.pvalue:.2e}")


def approve(income):
    """Stand-in for the model: approve when the score clears zero."""
    return (0.09 * income - 3.2 > 0).astype(int)


# Prediction drift. Measurable on day one, with no labels at all.
print()
print("approval rate at training time:", round(approve(reference).mean(), 3))
print("approval rate now            :", round(approve(live).mean(), 3))
Output
income  mean  39.8 ->  32.7   PSI 0.338   KS 0.244   p 8.62e-131
age     mean  38.0 ->  38.1   PSI 0.005   KS 0.023   p 1.56e-01

approval rate at training time: 0.639
approval rate now            : 0.398

Reading the four numbers

PSI 0.338 for income. The Population Stability Index compares how much of the data falls in each bin, then and now. The industry rule of thumb is: below 0.1 nothing happened, 0.1 to 0.25 worth a look, above 0.25 investigate now. Those thresholds are convention from credit scoring, not a theorem — treat them as a starting point you will tune.

PSI 0.005 for age. The control feature stayed quiet. A detector that fires on everything is a detector nobody reads.

KS 0.244 with p ≈ 1e-131. The Kolmogorov–Smirnov statistic is the largest gap between the two cumulative distributions, so 0.244 means the curves separate by about 24 percentage points at their widest. The p-value is astronomically small, and that is a warning rather than a result: with 5000 samples per side, a change far too small to matter would also produce a tiny p-value. Report the effect size, alert on the effect size. The p-value tells you the sample was large, which you already knew.

Approval rate 0.639 to 0.398. This is the line a business person will act on. No labels were needed, no statistics were needed, and it is visible the same day. Prediction drift is the cheapest monitor in existence and the one most often left unbuilt.

What a real monitoring set looks like

Five layers, cheapest first. Build them in this order, and notice that only the last one has to wait.

LayerExamplesAvailable
Operationalrequests per second, p99 latency, error rate, memoryimmediately
Data qualitynull rate per column, out-of-range counts, unseen categories, schema mismatchesimmediately
Predictionmean score, share of each predicted class, score histogramimmediately
Input driftPSI or KS per feature against a frozen referenceimmediately
Performanceaccuracy, precision, recall, calibrationwhen labels arrive

Data quality is the layer that earns its keep first. In practice, most "the model is drifting" investigations end at a renamed column or a unit change upstream.

Common mistakes

Alerting on the p-value. With production traffic volumes, every feature is "significantly different" every week. Alert on PSI, on a KS statistic, or on a relative change in the mean — quantities that stay flat when the change is small, whatever the sample size.

A reference window that slides. Compare against the frozen training distribution, not against last week. A slow drift compared to last week always looks like no drift, which is precisely how you miss the slow ones.

Monitoring features nobody uses. Sort your drift checks by feature importance. Drift in a feature the model barely weighs is a curiosity. Drift in the top feature is an incident.

Retraining on every alert. Retraining on drifted data with no labels teaches the model whatever a broken pipeline is currently emitting. Confirm the cause first.

Ignoring feedback loops. If your model decides who gets a loan, you only ever observe repayment for people it approved. Next year's training data is a filtered version of reality, shaped by last year's model. Reserve a small random slice of traffic that bypasses the model, so you keep some unbiased ground truth.

Storing predictions without the inputs. When you finally get labels, you need the exact feature values that produced each prediction, plus the model version. Log the request, the response and the version together, or the investigation is impossible.

Try it yourself

Change live to rng.normal(38, 12, 5000) — a shift of two thousand rupees instead of seven.

Look at what the two numbers do. The p-value stays tiny. The PSI drops to something small. That contrast, on your own screen, is the whole argument for alerting on effect size instead of significance.

What to learn next

Researcher — Mathematics and papers.

Naming the shifts precisely

Let the training distribution be $p_{\text{tr}}(x, y)$ and the deployment distribution $p_{\text{te}}(x, y)$. The joint factorises two ways, and each factorisation names a family:

  • Covariate shift — $p_{\text{tr}}(x) \neq p_{\text{te}}(x)$, with $p(y \mid x)$ unchanged. The population changed; the underlying rule did not.
  • Label shift (prior probability shift) — $p_{\text{tr}}(y) \neq p_{\text{te}}(y)$, with $p(x \mid y)$ unchanged. Natural where $y$ causes $x$, as in disease prevalence changing while symptoms do not.
  • Concept drift — $p(y \mid x)$ itself changes. No reweighting of inputs fixes this; the target function moved and you need new labels.

Only covariate shift is correctable without labels. Under covariate shift with shared support, importance weighting by $w(x) = p_{\text{te}}(x) / p_{\text{tr}}(x)$ gives an unbiased risk estimate (Shimodaira, 2000). Estimating $w$ well is itself hard, and the variance of the weighted estimator scales with the Rényi divergence between the distributions — so the correction degrades exactly when it is most needed.

Two-sample testing

Kolmogorov–Smirnov. $D_{n,m} = \sup_x |F_n(x) - G_m(x)|$, the maximum vertical gap between empirical CDFs. Univariate and continuous only, insensitive in the tails, and it detects a difference without telling you where.

Population Stability Index. With reference proportions $e_i$ and actual proportions $a_i$ over $B$ bins:

$$\mathrm{PSI} = \sum_{i=1}^{B} (a_i - e_i) \ln\frac{a_i}{e_i}$$

This is the symmetrised Kullback–Leibler divergence, $\mathrm{KL}(a | e) + \mathrm{KL}(e | a)$, also called Jeffreys divergence. It is a divergence, not a test — no p-value, no null distribution. The 0.1 and 0.25 thresholds are credit-industry convention with no distributional justification, and they depend on bin count: more bins mechanically increase PSI. Fix $B$ and compare like with like.

Maximum Mean Discrepancy. For a characteristic kernel $k$, $\mathrm{MMD}^2(P, Q) = |\mu_P - \mu_Q|^2_{\mathcal{H}}$ in the RKHS, with an unbiased estimator computable from the Gram matrices (Gretton et al., 2012). Multivariate, kernel-based, and a permutation test gives a valid p-value. Cost is $O(n^2)$ naively, with linear-time variants available. This is the principled choice when you need one test over the whole feature vector rather than $d$ marginal tests.

Classifier two-sample test. Train a discriminator to separate reference from live samples. Its held-out AUC is an interpretable effect size — 0.5 means indistinguishable — and the feature importances name the culprits. Lopez-Paz and Oquab (2017) formalise it.

Multiple testing

Running a KS test per feature per day on $d$ features is $d$ tests per day. At $d = 200$ and $\alpha = 0.05$, ten false alarms daily by construction. Apply Benjamini–Hochberg FDR control across features, or move to a single multivariate test.

Better still, decouple the two questions. Statistical significance answers "did anything change". Effect size answers "does it matter". Only the second belongs on an alert.

Label shift estimation without labels

Black Box Shift Estimation (Lipton, Wang and Smola, 2018) recovers the deployment label distribution from an existing classifier's confusion matrix on held-out training data plus its predicted-label distribution on unlabelled deployment data:

$$\hat{w} = \hat{C}^{-1} \hat{q}$$

where $\hat{C}$ is the confusion matrix (predicted against true, from held-out training data), $\hat{q}$ is the vector of predicted-label proportions in deployment, and $\hat{w}$ estimates the ratio $p_{\text{te}}(y) / p_{\text{tr}}(y)$. Requires $\hat{C}$ invertible and the label-shift assumption to hold. RLLS (Azizzadenesheli et al., 2019) adds regularisation for stability with few samples.

Streaming detectors

For sequential settings with fast labels:

  • DDM (Gama et al., 2004) monitors the online error rate, flagging warning and drift when it exceeds a running minimum by two and three standard deviations respectively.
  • ADWIN (Bifet and Gavaldà, 2007) maintains an adaptive window, splitting it whenever two sub-windows have significantly different means, with a rigorous false-positive bound.
  • Page–Hinkley — a cumulative-sum test for a shift in mean.

All three need labels promptly, which rules them out for the delayed-label problems that dominate industrial ML.

Performance estimation without labels

An active area, and one to treat sceptically. Confidence-based methods (CBPE) assume the model stays calibrated under shift, which is exactly what shift breaks. Importance-weighted estimation (DLE, and reverse-classifier approaches) is better founded but degrades as the distributions separate.

The honest summary: these methods are useful as early-warning indicators and unsound as replacements for measured performance. Design the label collection path instead of hoping to avoid needing it.

Papers

  • Gretton et al., A Kernel Two-Sample Test, JMLR 2012 — jmlr.org/papers/v13/gretton12a.html
  • Lipton, Wang and Smola, Detecting and Correcting for Label Shift with Black Box Predictors, ICML 2018 — arxiv.org/abs/1802.03916
  • Rabanser, Günnemann and Lipton, Failing Loudly: An Empirical Study of Methods for Detecting Dataset Shift, NeurIPS 2019 — arxiv.org/abs/1810.11953
  • Lopez-Paz and Oquab, Revisiting Classifier Two-Sample Tests, ICLR 2017 — arxiv.org/abs/1610.06545
  • Gama et al., Learning with Drift Detection, SBIA 2004
  • Bifet and Gavaldà, Learning from Time-Changing Data with Adaptive Windowing, SDM 2007
  • Quiñonero-Candela et al., Dataset Shift in Machine Learning, MIT Press 2009 — the book-length treatment.

What to learn next