Monitoring Models in Production

Monitoring when you have no labels

Long before a true outcome arrives, you can already watch the model's own predictions for warning signs, using signals that never need a label at all.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

You can watch a model's own predictions for trouble long before any true outcome arrives, using signals that never need a label.

The analogy you have already lived

A cook checking a pressure cooker cannot lift the lid to taste the food mid-cook. The lid stays sealed under pressure.

They still know a great deal without tasting. The whistle count. The smell escaping the valve. The colour of steam. None of that is the final taste test, and all of it is useful right now.

Monitoring a model without labels works the same way. You cannot taste the true outcome yet. You can still watch what is happening right in front of you.

Why it exists

The last lesson showed that true outcomes often arrive weeks or months late. Waiting for them to check on a model means waiting weeks or months to notice a problem.

That wait is too slow for anything urgent. A team needs signals it can check today, using only the inputs and the predictions the model already made.

How it works

   inputs  ---->  [ model ]  ---->  prediction (today, instantly)
                                          |
                          watch THIS for trouble, right now
                                          |
                                    (the true outcome
                                     is still weeks away)

Nothing here proves the model is right. It shows whether the model is behaving like itself, compared to how it behaved on the data it was trained on.

A real example you have seen

A loan app quietly ran a marketing push that brought in higher earners than usual. Within a day, far more applicants than normal were being approved.

Nobody had a single confirmed repayment outcome yet — those take months. The sudden jump in the approval rate itself was the first, immediate clue that something about the incoming applicants had changed.

The honest part

None of these signals can tell you the model is correct. A model can be confidently, smoothly wrong in a way that never trips a label-free check at all.

Treat these as a smoke alarm, not a taste test. They buy you time to look closer. They do not replace eventually checking against real outcomes.

Remember this

  • Predictions and inputs are available immediately; true outcomes are not.
  • Watch how predictions behave compared to training, as an early warning.
  • These checks catch "something changed", never "the model is right".

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy pandas scikit-learn

Two free signals: the approval rate, and confidence

Reusing the loan-scoring model from model serving, watch how its own predictions shift when a marketing push changes who applies — with no labels involved anywhere in this script.

no_labels.py
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

rng = np.random.RandomState(0)

# Same loan-scoring setup as the model-serving lesson.
n = 400
X = pd.DataFrame({
    "income": rng.uniform(5, 80, n).round(1),
    "years": rng.uniform(0, 10, n).round(1),
    "age": rng.randint(21, 65, n).astype(float),
})
score = 0.05 * X["income"] + 0.35 * X["years"] + 0.01 * X["age"] - 3.5
y = (score + rng.normal(0, 0.8, n) > 0).astype(int)
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000)).fit(X, y)

# ---- No labels exist yet for what happens next. Only inputs and predictions. ----

# Week 1: applicants who look like training data.
week1 = pd.DataFrame({
    "income": rng.uniform(5, 80, 500).round(1),
    "years": rng.uniform(0, 10, 500).round(1),
    "age": rng.randint(21, 65, 500).astype(float),
})

# Week 8: a marketing push started attracting higher earners.
week8 = pd.DataFrame({
    "income": rng.uniform(40, 120, 500).round(1),
    "years": rng.uniform(0, 10, 500).round(1),
    "age": rng.randint(21, 65, 500).astype(float),
})

p1 = model.predict_proba(week1)[:, 1]
p8 = model.predict_proba(week8)[:, 1]

print(f"week 1  -- predicted approval rate: {(p1 >= 0.5).mean():.1%}   mean confidence: {p1.mean():.3f}")
print(f"week 8  -- predicted approval rate: {(p8 >= 0.5).mean():.1%}   mean confidence: {p8.mean():.3f}")

low_conf1 = ((p1 > 0.4) & (p1 < 0.6)).mean()
low_conf8 = ((p8 > 0.4) & (p8 < 0.6)).mean()
print(f"week 1  -- fraction the model is unsure about (0.4-0.6): {low_conf1:.1%}")
print(f"week 8  -- fraction the model is unsure about (0.4-0.6): {low_conf8:.1%}")
Output
week 1  -- predicted approval rate: 69.4%   mean confidence: 0.683
week 8  -- predicted approval rate: 96.8%   mean confidence: 0.939
week 1  -- fraction the model is unsure about (0.4-0.6): 5.6%
week 8  -- fraction the model is unsure about (0.4-0.6): 2.4%

These are exact numbers from this seeded script. Not one true repayment outcome was used anywhere above. The 96.8% approval rate alone is reason enough to go check what changed.

Line-by-line walkthrough

week1 is drawn from the same income range the model trained on. week8 shifts that range upward, standing in for a real population change.

predict_proba(...)[:, 1] reads the model's own confidence for the positive class, 0.5 and above meaning "approve". Both weeks compute this identically; nothing about the check depends on the true outcome existing yet.

The "unsure" band (0.4 to 0.6) counts predictions near the model's own decision boundary. Watch what happened to it here: it went down, not up.

A trap worth naming out loud

You might expect the model to grow less confident on unfamiliar, higher-income applicants. It grew more confident instead.

A logistic regression model has no concept of "outside my training range". It applies the same straight-line rule everywhere, and rewards inputs further from the boundary with a more extreme, more confident score. Confidence going up is not proof of health. It can be the model cheerfully extrapolating into territory it has never seen.

Common mistakes

Trusting rising confidence as a sign of a healthy model. As shown above, confidence can rise for the wrong reason. Pair it with the input-drift check from measuring data drift, which would have flagged this income shift directly.

Watching only the overall approval rate. A stable overall rate can hide one segment approving far more and another far less. Monitoring by segment catches what an aggregate number hides.

Treating a label-free alert as a confirmed model failure. It is a prompt to investigate, not a verdict. Escalate proportionally, and confirm with real outcomes once they arrive.

Try it yourself

Compute the input PSI on income between week1 and week8, using the psi() function from measuring data drift. Check whether it agrees with the prediction-based signals above.

What to learn next

Researcher — Mathematics and papers.

The general class of label-free signals

These fall into three families, in increasing order of how much they assume about the model's internals:

  • Input-based — distributional checks on $P(x)$ directly, covered in measuring data drift. Model-agnostic, but blind to concept drift entirely.
  • Output-based — distributional checks on $P(\hat{y} \mid x)$, the model's own prediction distribution, as used above. Cheap, but only sensitive to shifts that actually move the model's decision function.
  • Internal-state-based — checks on intermediate representations (embedding-space distances, activation statistics), requiring access to the model's internals rather than only its API.

Why confidence is not a reliable proxy for correctness

A model trained with a standard cross-entropy or log-loss objective is calibrated only within the support of its training distribution, and calibration guarantees say nothing about behaviour outside that support. For a linear model, confidence is a monotonic function of distance from the decision boundary in feature space, which increases without bound as an input moves away from the training region — the mechanism behind the "confidence went up" result above.

This motivates uncertainty-aware alternatives that behave more conservatively out-of-distribution: Monte Carlo dropout, deep ensembles (Lakshminarayanan et al., 2017), and conformal prediction (Vovk et al., 2005), which produces prediction sets with a distribution-free coverage guarantee rather than a single point confidence.

A formal detection framing

Treat the reference (training-time) prediction distribution $P_{\text{ref}}(\hat{y})$ and the live prediction distribution $P_{\text{live}}(\hat{y})$ as two samples, and apply any two-sample test — PSI, KS, or MMD, all introduced in the data-drift lesson — directly to the model's output score instead of to an input feature. This is sometimes called prediction drift or output drift detection, and it is cheaper to compute than most input-drift checks because it requires monitoring only one number (or one small vector, for multi-class) per prediction, already produced as a side effect of serving.

Papers

  • Lakshminarayanan, Pritzel and Blundell, Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles, NeurIPS 2017 — arxiv.org/abs/1612.01474
  • Vovk, Gammerman and Shafer, Algorithmic Learning in a Random World, Springer 2005 — the foundational text on conformal prediction.
  • Guo, Pleiss, Sun and Weinberger, On Calibration of Modern Neural Networks, ICML 2017 — arxiv.org/abs/1706.04599

What to learn next