Monitoring Models in Production

Working with delayed labels

A label is the true outcome you are trying to predict, and in most real systems it arrives weeks after the prediction, which makes "how accurate are we today" a harder question than it sounds.

Read these first

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A delayed label is the true outcome of a prediction, arriving days or months late.

The analogy you have already lived

A farmer chooses which seeds to plant in June. Whether that choice was good is not known in June.

It is known months later, at harvest, when the crop either came in well or did not. The farmer cannot taste the outcome of a decision on the day it is made.

A model that scores loan applications has the exact same wait. It predicts today. Whether it was right is only known when the loan is repaid or defaults, months from now.

Why it exists

A label here means the real, confirmed answer. Did the loan actually get repaid? Did the parcel actually arrive on time? Did the transaction actually turn out to be fraud?

Most interesting predictions are about the future. "Will this loan be repaid" only has a true answer once the loan finishes, one way or another.

You cannot compute accuracy without the true answer. So for a long stretch after a prediction, accuracy for that prediction genuinely does not exist yet. It is not hidden. It has not happened.

How it works

   day 0            day 0 to ~180                day 180
   prediction  ---->  outcome is unknown  ---->  label finally arrives
   made now           to anyone, including            (repaid or defaulted)
                       the monitoring system

Anything you measure using "today's" labels is secretly only measuring the fast-resolving cases. The slow ones have not reported in yet.

A real example you have seen

Credit card fraud is confirmed fast, sometimes within hours, once a cardholder disputes a charge. A missed EMI on a loan is confirmed slowly, sometimes only after 90 days of non-payment.

A monitoring dashboard checked today mixes fully-confirmed fraud cases with loan cases still waiting to reveal themselves. Comparing the two numbers directly is comparing different things.

The honest part

This one surprises almost everyone the first time. A naive "accuracy today" number is not only noisy while labels trickle in. It can be systematically wrong, in one fixed direction, not only imprecise.

Cases that resolve fast are not a random sample of all cases. They tend to be the extreme ones. Building any metric on labels-so-far quietly measures the wrong population.

Remember this

  • A label is the confirmed true outcome, and it often arrives long after the prediction.
  • Metrics built on whatever labels have arrived so far are not a random sample of outcomes.
  • Fast-arriving labels are usually the unusual cases, which biases any "so far" number.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy pandas

Seeing the bias, not only describing it

Bad loan outcomes reveal themselves fast, within a month. Good outcomes are only confirmed when the loan matures, much later. Snapshot the data "today" and compare what you'd naively compute against the truth.

delayed_labels.py
import numpy as np
import pandas as pd

rng = np.random.RandomState(0)

n = 2000
today = 200

issue_day = rng.uniform(0, today, n)
is_bad = rng.binomial(1, 0.20, n)  # true default rate is 20%

# Bad outcomes (a default) reveal themselves fast: 5-30 days.
# Good outcomes (fully repaid) only get confirmed at loan maturity: 180 days.
delay = np.where(is_bad == 1, rng.uniform(5, 30, n), 180)
label_ready_day = issue_day + delay
observed = label_ready_day <= today

df = pd.DataFrame({"issue_day": issue_day, "is_bad": is_bad, "observed": observed})

true_bad_rate = df["is_bad"].mean()
naive_bad_rate = df.loc[df["observed"], "is_bad"].mean()

print(f"loans issued:               {n}")
print(f"labels observed by today:   {df['observed'].sum()} ({df['observed'].mean():.1%})")
print(f"TRUE bad rate (all loans):        {true_bad_rate:.1%}")
print(f"NAIVE bad rate (observed only):   {naive_bad_rate:.1%}")
Output
loans issued:               2000
labels observed by today:   516 (25.8%)
TRUE bad rate (all loans):        19.5%
NAIVE bad rate (observed only):   70.5%

These are exact numbers from this seeded script. The true bad rate is 19.5%. Looking only at labels that have arrived so far reports 70.5% — more than three times too high, from real, correctly-computed data.

Line-by-line walkthrough

is_bad is drawn once, honestly, at a true 20% rate. It never changes. This is the ground truth the monitoring system is trying to estimate.

delay gives bad outcomes a short, fast window and good outcomes a long, fixed one. This single line is the entire mechanism behind the biased number below it.

observed is True only for loans whose label has actually arrived by day 200. A monitoring job querying "today's known outcomes" sees exactly this subset, nothing more.

The gap between the two printed rates is not a bug in the code. It is what naive "so far" monitoring looks like on labels with unequal delay, computed correctly.

Common mistakes

Reporting accuracy-so-far without saying how much data has actually resolved. "516 of 2000 labels" changes how anyone should read "70.5%". Always publish the observed fraction alongside the metric.

Comparing this week's naive rate against last week's naive rate. Both are biased the same structural way, so the comparison can look stable while the true rate is moving. Compare fully-matured cohorts against each other instead.

Assuming the delay is the same for every outcome. The bias in this lesson exists specifically because delay differs by outcome. If delay were random and equal for every case, the naive rate would be noisy but not systematically wrong.

Backfilling metrics without ever revisiting them. A dashboard showing "this month's accuracy" needs a second, later pass once that month's labels have matured, or it permanently reports the biased early read.

Try it yourself

Change the good-outcome delay from a fixed 180 to rng.uniform(150, 210, n), so maturity varies a little. Rerun, and check whether the naive-versus-true gap shrinks, stays the same, or grows.

What to learn next

Researcher — Mathematics and papers.

This is informative censoring

The statistics literature has a name for exactly this pattern: informative right-censoring. An outcome's censoring status (whether it has been observed yet) is not independent of the outcome's value, which breaks the assumption behind treating "observed so far" as a random sample.

This is distinct from missing completely at random (MCAR), where any naive estimator computed on observed cases only remains unbiased. Here it is neither MCAR nor even MAR (missing at random conditional on covariates) in general, because the missingness mechanism depends on the unobserved outcome itself.

Correcting for it

Vintage / cohort analysis. Group by issue date, and only report a cohort's rate once its full observation window has elapsed. This sidesteps the bias by refusing to compute a metric before it can be computed honestly.

Survival analysis. Model $P(\text{bad} \mid \text{not yet observed at time } t)$ directly using a Kaplan-Meier estimator or a Cox proportional hazards model, rather than treating right-censored cases as missing. This recovers a corrected rate from partial information, instead of discarding it.

$$\hat{S}(t) = \prod_{t_i \le t} \left(1 - \frac{d_i}{n_i}\right)$$

Where $\hat{S}(t)$ is the estimated probability of remaining "good" past time $t$, $d_i$ is the number of bad outcomes observed at time $t_i$, and $n_i$ is the number still at risk (not yet resolved either way) immediately before $t_i$.

Proxy labels. Use a faster, noisier signal correlated with the true outcome — an early delinquency flag standing in for eventual default — accepting some label noise in exchange for a shorter delay. Monitoring when you have no labels covers the case where no proxy is available at all.

Complexity and cost

Kaplan-Meier estimation over $n$ events is $O(n \log n)$, dominated by sorting event times. Cohort analysis is $O(n)$ per cohort but trades that simplicity for a longer wait before any number is trustworthy — the correction is procedural, not computational.

Papers

  • Kaplan and Meier, Nonparametric Estimation from Incomplete Observations, JASA 1958 — the origin of the estimator above.
  • Rubin, Inference and Missing Data, Biometrika 1976 — the MCAR / MAR / MNAR framework this problem sits inside.
  • Cox, Regression Models and Life-Tables, JRSS-B 1972 — the proportional hazards model used for covariate-adjusted survival estimates.

What to learn next