Shipping Vision Models

Detecting drift in image data

Cameras get dusty, lighting changes and new things walk into frame, so a vision model degrades silently unless something is watching the pictures themselves.

On this page 11
  1. The short answer
  2. The analogy you have lived
  3. Why nothing warns you
  4. What you can watch instead
  5. The test that explains itself best
  6. The trap that catches every team
  7. The one that is hardest to catch
  8. Where you have seen this
  9. The honest part
  10. Remember this
  11. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Drift means the pictures arriving today are not like the ones the model learned from. Nothing will tell you unless you look.

The analogy you have lived

A camera goes up on a wall outside a building. For the first few months it works beautifully.

Two years later it is still working, and it is much worse. Dust has settled on the lens. The streetlight outside was replaced with a brighter white one. A tree in the corner of the frame has grown into it. The wall behind was repainted.

No single day changed anything. Nobody made a decision. The model was never touched. And it is now looking at a different world.

Why nothing warns you

The model still runs. It still returns an answer, with a confidence, in the usual number of milliseconds. Your dashboards are green.

Only one thing would reveal the problem: knowing whether those answers are right. In production, nobody tells you. A shop counting visitors does not get corrected. A defect detector does not receive a report on the defects it missed.

So the honest question becomes: can you notice the world changing, without knowing whether you are still right?

What you can watch instead

Three things, in increasing order of usefulness.

The pictures themselves. Average brightness, contrast, sharpness, colour balance, size. Cheap to compute, and they catch the physical failures — a dirty lens, a moved camera, a changed light.

The model's internal description of each picture. Before the final answer, a model builds a summary of what it saw. Comparing today's summaries against a stored reference catches changes that simple brightness statistics miss.

The model's answers. If the fraction of pictures called "defective" doubles overnight, something happened. Either the world changed or the model did.

The test that explains itself best

Here is a trick worth knowing. Take a few hundred pictures from when the model was working, and a few hundred from today. Mix them. Now train a small, simple classifier to guess which pile each picture came from.

If it cannot do better than a coin toss, the two piles are alike. No drift.

If it guesses correctly nearly every time, the two piles are plainly different. You can then go and look at what the classifier is picking up on.

The score of that little classifier is a drift measurement anyone can understand.

The trap that catches every team

Statistical tests get more sensitive as you feed them more data. That sounds good and it is a problem.

With a few hundred pictures, a test will flag only real changes. With twenty thousand, the same test flags changes so small that the model does not notice them.

So an alarm on its own means very little. What matters is how big the change is, not whether a test noticed it. A tiny, certain change is not worth waking anyone.

The one that is hardest to catch

Here is the uncomfortable finding. The easiest change to detect is a boring one — the sun moved, everything got brighter.

The most dangerous change is the hardest to detect: something genuinely new walking into frame. It affects a small fraction of pictures, so the overall statistics barely move.

The change you most want to catch is the one your monitoring is worst at. That is worth knowing before you trust a dashboard.

Where you have seen this

  • A face unlock that stopped working after you grew a beard.
  • A number-plate camera that got worse after the road was resurfaced and reflections changed.
  • A shop counter that started over-counting when a mirror was installed.
  • A crop-monitoring model whose accuracy fell after a season with unusual rain.

The honest part

Drift detection tells you something changed. It does not tell you your model got worse, and it does not tell you what to do.

The only real answer is labels. Sample production images every month, have a person label them, and measure. Everything else is an early warning that tells you when to go and look.

Remember this

  • Cameras and the world change slowly and silently. Nothing errors.
  • The easiest test to explain: can a small classifier tell old pictures from new ones?
  • A statistical alarm on a large sample means very little. Size of change beats certainty of change.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scipy scikit-learn

Three detectors, four scenarios, and the one experiment that stops you from building an alarm nobody trusts.

Four kinds of drift, three ways to look for them

image_drift.py
import numpy as np
from scipy.stats import ks_2samp
from scipy.ndimage import gaussian_filter
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import roc_auc_score

rng = np.random.default_rng(0)
S = 32


def photos(n, brightness=0.0, blur=0.0, extra_class=False):
    """Fake camera frames: a bright shape on a dark background."""
    out = np.full((n, S, S), 0.15, np.float32)
    gy, gx = np.mgrid[0:S, 0:S]
    for i in range(n):
        kind = rng.integers(0, 3 + int(extra_class))
        r = rng.integers(6, 11)
        cy, cx = rng.integers(r + 1, S - r - 1, 2)
        if kind == 0:
            m = (gy - cy) ** 2 + (gx - cx) ** 2 <= r * r
        elif kind == 1:
            m = (abs(gy - cy) <= r) & (abs(gx - cx) <= r)
        elif kind == 2:
            m = (gy - cy >= -r) & (gy - cy <= r) & (abs(gx - cx) <= (gy - cy + r) / 2)
        else:
            m = (abs(gy - cy) + abs(gx - cx)) <= r      # a shape never seen in training
        out[i][m] = 0.85
        out[i] += rng.normal(0, 0.02, (S, S))
        if blur:
            out[i] = gaussian_filter(out[i], blur)
    return np.clip(out + brightness, 0, 1)


# A stand-in for a model's penultimate features: a fixed random projection of
# the image, which is deterministic and needs no training.
PROJ = rng.normal(size=(S * S, 32)) / np.sqrt(S * S)


def embed(x):
    return x.reshape(len(x), -1) @ PROJ


def ks_drift(a, b, alpha=0.05):
    p = np.array([ks_2samp(a[:, d], b[:, d]).pvalue for d in range(a.shape[1])])
    return int((p < alpha / a.shape[1]).sum()), float(p.min())   # Bonferroni


def mmd_rbf(a, b, gamma=None, perms=200):
    z = np.vstack([a, b])
    d2 = ((z[:, None] - z[None]) ** 2).sum(-1)
    gamma = gamma or 1.0 / np.median(d2[d2 > 0])
    K = np.exp(-gamma * d2)
    n = len(a)

    def stat(idx):
        i, j = idx[:n], idx[n:]
        return (K[np.ix_(i, i)].mean() + K[np.ix_(j, j)].mean()
                - 2 * K[np.ix_(i, j)].mean())

    order = np.arange(len(z))
    obs = stat(order)
    null = [stat(rng.permutation(order)) for _ in range(perms)]
    return float(obs), float((np.sum(np.array(null) >= obs) + 1) / (perms + 1))


def classifier_auc(a, b):
    X = np.vstack([a, b])
    y = np.r_[np.zeros(len(a)), np.ones(len(b))]
    xtr, xte, ytr, yte = train_test_split(X, y, test_size=.4, random_state=0,
                                          stratify=y)
    clf = LogisticRegression(max_iter=2000).fit(xtr, ytr)
    return float(roc_auc_score(yte, clf.predict_proba(xte)[:, 1]))


ref = embed(photos(500))

CASES = {
    "no drift (same camera)":       photos(500),
    "brightness +0.15 (sun moved)": photos(500, brightness=0.15),
    "lens gone slightly soft":      photos(500, blur=1.2),
    "a new shape appears":          photos(500, extra_class=True),
}

print(f"{'scenario':<30} {'KS dims':>8} {'min p':>10} {'MMD p':>7} {'clf AUC':>8}")
for name, imgs in CASES.items():
    cur = embed(imgs)
    n_sig, pmin = ks_drift(ref, cur)
    _, pm = mmd_rbf(ref[:200], cur[:200])
    print(f"{name:<30} {n_sig:>4}/32 {pmin:>10.2e} {pm:>7.3f} "
          f"{classifier_auc(ref, cur):>8.3f}")

print("\ntruly no drift, at three sample sizes:")
for n in (200, 2000, 20000):
    a, b = embed(photos(n)), embed(photos(n))
    n_sig, pmin = ks_drift(a, b)
    print(f"  n = {n:>6}: {n_sig:>2}/32 flagged, smallest p = {pmin:.2e}, "
          f"AUC {classifier_auc(a, b):.3f}")

print("\na harmless +0.01 brightness shift, at the same sample sizes:")
for n in (200, 2000, 20000):
    a, b = embed(photos(n)), embed(photos(n, brightness=0.01))
    n_sig, pmin = ks_drift(a, b)
    print(f"  n = {n:>6}: {n_sig:>2}/32 flagged, smallest p = {pmin:.2e}, "
          f"AUC {classifier_auc(a, b):.3f}")
Output
scenario                        KS dims      min p   MMD p  clf AUC
no drift (same camera)            0/32   1.34e-02   0.761    0.500
brightness +0.15 (sun moved)     20/32   9.37e-56   0.005    0.968
lens gone slightly soft           6/32   1.00e-05   0.005    0.573
a new shape appears               3/32   2.48e-04   0.035    0.553

truly no drift, at three sample sizes:
  n =    200:  0/32 flagged, smallest p = 5.21e-02, AUC 0.497
  n =   2000:  0/32 flagged, smallest p = 1.81e-02, AUC 0.529
  n =  20000:  0/32 flagged, smallest p = 4.64e-02, AUC 0.495

a harmless +0.01 brightness shift, at the same sample sizes:
  n =    200:  0/32 flagged, smallest p = 2.97e-02, AUC 0.543
  n =   2000:  2/32 flagged, smallest p = 7.12e-04, AUC 0.548
  n =  20000: 16/32 flagged, smallest p = 3.47e-15, AUC 0.567

The main table, read row by row

No drift gives an AUC of exactly 0.500 and flags nothing. All three detectors are calibrated. That row is the control, and any monitoring system that cannot produce it is broken before it starts.

A brightness shift is easy to detect: 20 of 32 dimensions flagged, a p-value of 9.37e-56, AUC 0.968. A global lighting change moves every feature. This is the drift your monitoring will catch, and it is the drift least likely to be the one that hurts you.

A softened lens is much subtler: 6 dimensions flagged, AUC 0.573. Detected, but barely. Notice how differently the two views read: the p-value of 1.00e-05 sounds decisive while the AUC of 0.573 says the two sets are barely separable. The p-value is answering "is there any difference at all", the AUC "how different are they". Only the second is a business question.

And here is the uncomfortable one. A new, unseen shape entering the stream flags 3 of 32 dimensions with an AUC of 0.553 — the weakest signal in the table. The new class is a quarter of the images, so it moves the pooled statistics only slightly. It is also, by a distance, the most dangerous of the four: the model has no class for it and will confidently assign it to something wrong.

The change you most want to catch is the one aggregate drift detection is worst at. Take that seriously when you design alerts. Detecting a new class needs a per-image out-of-distribution score, not a batch-level test.

The sample-size experiment, which is the point of the lesson

With genuinely no drift, nothing is flagged at any sample size. 0 out of 32 at n=200, at n=2,000 and at n=20,000, with AUC hovering around 0.50. Bonferroni correction is doing its job and there are no false alarms. A test that fires on the null is a test you will learn to ignore.

With a harmless +0.01 brightness shift, the picture changes completely. At n=200, nothing is flagged. At n=2,000, two dimensions. At n=20,000, sixteen dimensions with a p-value of 3.47e-15.

Meanwhile the classifier AUC crawls from 0.543 to 0.567. The two distributions are nearly identical the whole time. A brightness shift of one percent is not going to change a single prediction.

So the same real-world change produces "all clear" or "sixteen alarms" depending only on how many images you looked at. Any alerting rule built on p-values will fire constantly once your traffic grows, and your team will disable it within a month.

The fix is to alert on effect size and use significance only as a gate:

python
def should_alert(ref, cur):
    n_sig, _ = ks_drift(ref, cur)
    auc = classifier_auc(ref, cur)
    return n_sig > 0 and auc > 0.75      # certain AND large

Pick the AUC threshold by measuring it against known-harmless changes on your own data. On the table above, 0.75 would fire on the brightness shift and stay quiet on everything else — which is exactly right if a brightness shift is what breaks your model, and exactly wrong if new classes are. Tune it against your failure, not against a rule of thumb.

What to compute in production

Layer the checks by cost.

python
def cheap_stats(img):                    # per image, microseconds
    return {"mean": img.mean(), "std": img.std(),
            "p01": np.percentile(img, 1), "p99": np.percentile(img, 99),
            "sharpness": np.abs(np.diff(img, axis=0)).mean(),
            "h": img.shape[0], "w": img.shape[1]}

Log those for every request. They catch a dirty lens, a moved camera and a changed encoder, and they cost nothing.

Then, on a schedule, run the batch tests above on the model's embeddings against a stored reference window. And separately track prediction statistics: the class distribution, the mean confidence, and the fraction of requests falling below your decision threshold.

Prediction drift deserves its own note. It is the cheapest signal of all and it is ambiguous by construction — the class mix changing is either the world changing or the model breaking, and the test cannot tell you which. Use it to trigger a look, never to conclude.

Common mistakes

Comparing against the training set for ever. The reference should be a window from when the model was known to be performing well, and it should be reviewed when you retrain.

Testing raw pixels instead of embeddings. Pixel statistics catch physical camera problems and nothing semantic. Both layers are needed; neither replaces the other.

One test per feature dimension without correction. Thirty-two independent tests at the 5 percent level give roughly an 80 percent chance of at least one false alarm. The alpha / a.shape[1] in the code is Bonferroni; a false-discovery-rate procedure is less conservative and usually better.

Alerting on p-values. Demonstrated above. It does not survive contact with production traffic.

Assuming drift means the model got worse. It means the input changed. Whether that matters is an empirical question, answered only by labels.

Skipping the label sample. Every drift signal is a proxy. Label a few hundred production images every month. It is the only measurement that is not a proxy, and it is the one teams cut first.

Try it yourself

Change the new-shape scenario so the new class is 90 percent of the images instead of a quarter, and watch the AUC climb. That tells you the detection threshold in terms of contamination rate — the fraction of unfamiliar images you would need before your monitoring notices. Then measure your production traffic against that number and decide whether it is good enough.

What to learn next

Researcher — Mathematics and papers.

What is being tested

Let $P_{\text{ref}}$ and $P_{\text{cur}}$ be distributions over inputs $x$, with labels $y$. The decomposition that matters:

  • Covariate shift: $P(x)$ changes, $P(y \mid x)$ fixed. Lighting, optics, sensor. Detectable from inputs alone.
  • Prior shift: $P(y)$ changes, $P(x \mid y)$ fixed. Seasonal mix, new site.
  • Concept drift: $P(y \mid x)$ changes. The definition of the label moved — a new defect category, a changed inspection standard.

Only the first two are detectable without labels. Concept drift is invisible to every method below, which is the fundamental limitation of unsupervised monitoring and the reason a labelled sample is not optional.

Maximum mean discrepancy

For a characteristic kernel $k$ with feature map $\phi$,

$$ \mathrm{MMD}^2(P, Q) = \left\lVert \mathbb{E}{x \sim P}[\phi(x)] - \mathbb{E}{y \sim Q}[\phi(y)] \right\rVert_{\mathcal{H}}^2 $$

with the unbiased estimator

$$ \widehat{\mathrm{MMD}}^2 = \frac{1}{m(m-1)}\sum_{i \ne j} k(x_i, x_j) + \frac{1}{n(n-1)}\sum_{i \ne j} k(y_i, y_j) - \frac{2}{mn}\sum_{i,j} k(x_i, y_j) $$

Gretton et al. (2012) established this as a consistent two-sample test: for a characteristic kernel, $\mathrm{MMD}(P,Q) = 0$ if and only if $P = Q$. The null distribution has no closed form, so significance comes from a permutation test, as in the code above. The RBF bandwidth is conventionally set by the median heuristic — $\gamma = 1/\text{median}(\lVert x_i - x_j\rVert^2)$ — which is a heuristic, not a principled choice, and Sutherland et al. (2017) show that learning the kernel gives substantially more power.

Cost is $O((m+n)^2 d)$ per permutation, which is why the example subsamples to 200 points per side.

The classifier two-sample test

Lopez-Paz and Oquab (2017) formalised the intuition: train a classifier to separate the two samples and test whether its held-out accuracy exceeds chance. Under $H_0: P = Q$, accuracy is binomially distributed around $1/2$, giving an exact test.

Its two practical advantages are decisive for production monitoring. It yields an interpretable effect size — an AUC of 0.57 and one of 0.97 mean different things, where two p-values below 1e-5 do not. And it is inspectable: the classifier's coefficients or feature importances point at what changed, which is the first question anyone asks after an alert.

Its weakness is finite-sample power: with a weak classifier or few samples it misses shifts a kernel test would catch. Use both.

The multiple-testing and power problem

Running $d$ marginal tests at level $\alpha$ gives a family-wise error rate of $1 - (1-\alpha)^d$ under independence — about 0.80 for $d = 32$, $\alpha = 0.05$. Bonferroni ($\alpha/d$) controls it conservatively; Benjamini-Hochberg controls the false discovery rate with more power and is the better default for monitoring.

The deeper issue is the one the sample-size experiment demonstrates. For a fixed true difference $\delta > 0$, the power of any consistent test tends to 1 as $n \to \infty$. Since $\delta$ is never exactly zero in production, every drift test fires eventually, given enough traffic. The measured progression — 0, then 2, then 16 dimensions flagged for a shift that moves AUC from 0.543 to 0.567 — is this asymptotic behaviour in miniature.

The correct construction is therefore an equivalence test, not a difference test: define a tolerance region and alert when the estimated effect size, with its confidence interval, leaves it. Effect size first, significance as a gate.

Detecting novelty, which is a different problem

The runnable table shows aggregate tests performing worst on the new-class scenario. That is structural: batch tests measure a difference between distributions, and a novel class contaminating a fraction $\epsilon$ of the stream moves aggregate statistics by roughly $O(\epsilon)$.

Per-image out-of-distribution scoring is the right tool:

  • Maximum softmax probability (Hendrycks and Gimpel, 2017) — the baseline; weak, because networks are overconfident on unfamiliar inputs.
  • ODIN (Liang et al., 2018) — temperature scaling plus input perturbation.
  • Mahalanobis distance (Lee et al., 2018) — class-conditional Gaussian fit to penultimate features; strong and cheap.
  • Energy score (Liu et al., 2020) — $-T\log\sum_c e^{f_c(x)/T}$; better calibrated than the softmax maximum.
  • kNN distance in feature space (Sun et al., 2022) — non-parametric, competitive, and needs a stored reference set.

Yang et al. (2022), OpenOOD, benchmark these consistently and are worth reading before choosing. Run one of them per image, and aggregate the fraction of images exceeding a threshold as your novelty signal — a rate, not a test.

Practical architecture

Reference window. A fixed sample from a period of verified good performance, stored as embeddings rather than images. Version it with the model, and refresh it when you retrain.

Two cadences. Cheap per-image statistics logged continuously; expensive batch tests on a schedule with a fixed window size, so the sample-size effect is held constant. A varying window size makes your alarm threshold meaningless.

Feature choice. Penultimate-layer embeddings are the standard, and they inherit the model's blind spots — a model that ignores a change will show no drift in its features even when the change matters. Pairing them with raw-pixel statistics covers different failures.

Labels. Sample, label, measure, monthly. Everything above is an early-warning system for deciding when to look; only labels close the loop. See monitoring and drift for the surrounding practice.

Tooling

alibi-detect implements MMD, the classifier test, learned-kernel MMD and several OOD detectors with a consistent interface. evidently and deepchecks target tabular data primarily and are weaker on images. torchdrift is PyTorch-native and lighter. For most teams the forty lines above, run on stored embeddings, is sufficient and easier to reason about than a framework.

Papers

  • Gretton et al., A Kernel Two-Sample Test, JMLR 2012
  • Lopez-Paz and Oquab, Revisiting Classifier Two-Sample Tests, ICLR 2017 — arxiv.org/abs/1610.06545
  • Rabanser et al., Failing Loudly: An Empirical Study of Methods for Detecting Dataset Shift, NeurIPS 2019 — arxiv.org/abs/1810.11953
  • Lee et al., A Simple Unified Framework for Detecting Out-of-Distribution Samples, NeurIPS 2018 — arxiv.org/abs/1807.03888
  • Liu et al., Energy-based Out-of-distribution Detection, NeurIPS 2020 — arxiv.org/abs/2010.03759
  • Yang et al., OpenOOD: Benchmarking Generalized Out-of-Distribution Detection, NeurIPS 2022 — arxiv.org/abs/2210.07242
  • Sutherland et al., Generative Models and Model Criticism via Optimized MMD, ICLR 2017 — arxiv.org/abs/1611.04488

What to learn next