Monitoring Models in Production
Monitoring by segment
Segment-level monitoring means checking metrics separately for meaningful slices of traffic, because an overall average can look calm while one slice is badly broken.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Segment-level monitoring means checking a metric separately for meaningful slices of traffic, instead of only one overall number.
The analogy you have already lived
A teacher announces the class average after an exam: seventy percent, a healthy score. Nobody is worried.
Hidden inside that one number: the front row scored ninety, and the back row scored fifty. The back row is struggling badly, and the healthy average hid it completely.
A single dashboard number for a whole model works exactly the same way. It can look fine while one group underneath it is not fine at all.
Why it exists
Traffic reaching a model is never one uniform thing. Different cities, different app versions, different sign-up channels, different device types.
A problem that hits one of those groups hard can be a small fraction of total volume. Averaged into everything else, it barely moves the overall number, even while it is a real, serious problem for that group.
How it works
OVERALL: 66.2% approval rate -- looks completely normal
broken down:
existing users 62.6% approval (4500 predictions)
new partner 98.4% approval (500 predictions) <-- this is the storyThe overall number is not wrong. It is the wrong resolution for noticing a problem that lives inside one slice.
A real example you have seen
A ride-hailing app's fare-estimate model can look accurate on average across a whole city. Split by neighbourhood, one area with poor GPS coverage is consistently wrong, every single day. A citywide average hides it completely.
The honest part
You cannot monitor every possible slice. There are too many combinations, and most of them do not have enough traffic to say anything reliable.
Pick the slices that matter for business or fairness reasons — channel, region, device, a known risky group — and monitor those on purpose. Treat "check everything" as a search you run occasionally, not a dashboard you maintain forever.
Remember this
- An overall metric can hide a serious problem confined to one segment.
- Choose segments that matter, rather than trying to monitor every possible slice.
- A small segment can be badly broken while barely moving the overall number.
What to learn next
- Alerts people do not ignore — turning a segment check into something that pages someone.
- When model metrics and business metrics disagree — another way an average number hides the real story.
- Concept drift — drift that hits one segment is often invisible in an overall number.
Developer — Code and libraries.
Setup
pip install numpy pandasThe same predictions, sliced two ways
Reusing the logged-prediction shape from logging every prediction, compare one overall approval rate against the same data split by channel.
import numpy as np
import pandas as pd
rng = np.random.RandomState(0)
# 4500 predictions from the existing, healthy channel.
existing = pd.DataFrame({
"channel": "existing_app",
"predicted_approve": rng.binomial(1, 0.62, 4500),
})
# 500 predictions from a brand-new partner channel, launched this week.
# Something about its applicants or its integration is producing far
# more approvals than normal -- but it is a small slice of total volume.
new_channel = pd.DataFrame({
"channel": "partner_api",
"predicted_approve": rng.binomial(1, 0.97, 500),
})
log = pd.concat([existing, new_channel], ignore_index=True)
overall_rate = log["predicted_approve"].mean()
print(f"OVERALL approval rate (all {len(log)} predictions): {overall_rate:.1%}")
print("\nBy segment:")
by_segment = log.groupby("channel")["predicted_approve"].agg(["mean", "count"])
for channel, row in by_segment.iterrows():
print(f" {channel:<14} approval rate: {row['mean']:.1%} n={int(row['count'])}")OVERALL approval rate (all 5000 predictions): 66.2% By segment: existing_app approval rate: 62.6% n=4500 partner_api approval rate: 98.4% n=500
Exact output from this seeded script. The overall number moved only four points above the healthy baseline. The segment underneath it moved thirty-six.
Line-by-line walkthrough
existing and new_channel are generated with different true approval rates (0.62 versus 0.97), standing in for two real populations behaving differently.
pd.concat merges them into one log, exactly as a real prediction table would hold both channels mixed together, arrival order interleaved.
groupby("channel") is the entire fix. Everything else in this script is the setup for why that one line matters.
Common mistakes
Choosing segments after seeing a problem, instead of in advance. By the time you know which slice to check, the damage from the previous outage in that slice is already done. Decide your standing segments before you need them.
Splitting so finely that each segment has too little data. A segment with twenty predictions produces a noisy rate that swings wildly by chance. Set a minimum sample size before trusting a segment's number, and group tiny segments into "other".
Only splitting by one dimension at a time. A problem can live at the intersection of two dimensions — a specific device type in a specific region — and be invisible in either dimension checked alone.
Re-deriving segment membership at query time from raw fields. Log the segment label itself at prediction time, alongside the prediction. It is one column, and it removes an entire category of "which definition of region did we use" arguments later.
Try it yourself
Add a third channel, "beta_users", with 50 predictions at a 0.30 true rate. Rerun, and check how little the overall number moves compared to how alarming that segment's own rate is.
What to learn next
- Alerts people do not ignore — turning a segment check into something that pages someone.
- When model metrics and business metrics disagree — another way an average number hides the real story.
- Concept drift — drift that hits one segment is often invisible in an overall number.
Researcher — Mathematics and papers.
Why aggregate stability is not a guarantee
For $k$ segments with sizes $n_1, \dots, n_k$ and rates $r_1, \dots, r_k$, the overall rate is the size-weighted mean:
$$r_{\text{overall}} = \frac{\sum_{i=1}^{k} n_i r_i}{\sum_{i=1}^{k} n_i}$$
A segment with small $n_i$ can have an arbitrarily extreme $r_i$ while contributing an arbitrarily small amount to $r_{\text{overall}}$. This is not a statistical subtlety, only arithmetic — but it is exactly why an aggregate dashboard is structurally unable to guarantee segment-level health.
Simpson's paradox, the sharper version of this problem
A stricter failure mode than "a segment is hidden": segment-level and aggregate trends can point in opposite directions. A model can improve within every single segment individually, while the overall metric goes down, if the mix of segment sizes shifts at the same time. Any monitoring system that reports a global trend without also reporting mix shift is vulnerable to reading this backwards.
Choosing which segments to monitor systematically
Rather than hand-picking segments, slice discovery methods search combinations of categorical attributes for subgroups with unusually large error or metric deviation, correcting for the multiple-comparisons problem inherent in testing many candidate slices:
- Chung et al. (2019), Slice Finder, formalises this as a search over a lattice of feature-value combinations, using an effect-size and statistical-significance filter to surface actionable slices instead of every combination.
- Sagadeeva and Boehm (2021), SliceLine, reformulates the same search as a matrix-based optimisation for tractable computation over large feature spaces.
Multiple comparisons
Monitoring $m$ segments simultaneously, each with its own significance test at level $\alpha$, inflates the family-wise false-alarm rate to approximately $1-(1-\alpha)^m$ under independence. A Bonferroni correction ($\alpha/m$ per test) or a false-discovery-rate procedure (Benjamini-Hochberg, 1995) is the standard correction once the segment count grows past a handful.
Papers
- Chung, Dae, Karlaš, Zhang and Krishnan, Slice Finder: Automated Data Slicing for Model Validation, ICDE 2019 — arxiv.org/abs/1807.06068
- Sagadeeva and Boehm, SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model Debugging, SIGMOD 2021
- Benjamini and Hochberg, Controlling the False Discovery Rate, JRSS-B 1995
What to learn next
- Alerts people do not ignore — turning a segment check into something that pages someone.
- When model metrics and business metrics disagree — another way an average number hides the real story.
- Concept drift — drift that hits one segment is often invisible in an overall number.