Monitoring Models in Production

When model metrics and business metrics disagree

A model metric measures how well the model does its narrow job, and a business metric measures whether that job actually helped, and the two can move in opposite directions.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A model metric measures how well the model does its narrow task. A business metric measures whether that helped at all. The two can disagree.

The analogy you have already lived

A delivery worker is judged on deliveries completed per hour. To hit a higher number, they start leaving packages at the gate without confirming anyone is home.

Their tracked number climbs. Complaints climb with it, and so do returned or stolen packages. The measured thing went up. The thing anyone actually cared about went down.

A model can do the same thing to a metric picked for it, without anyone intending it to.

Why it exists

A model is trained to get good at one narrow, measurable thing. Predicting a click. Predicting a rating. Predicting whether a loan gets repaid.

The business rarely cares about that narrow thing directly. It cares about revenue, retention, safety, or trust. Those are harder to measure at prediction time, and often only visible weeks later.

The narrow thing and the real thing are rarely perfectly aligned. A model can get better at the narrow thing while quietly making the real thing worse.

How it works

   MODEL METRIC: "did they click?"        BUSINESS METRIC: "did we make money?"
        goes UP  ------------->                    goes DOWN
                                     |
              the model learned to show cheap, clicky items
                over solid, higher-value ones

Both numbers are real and correctly measured. They are not measuring the same thing.

A real example you have seen

A shopping app's "you might like" section can be tuned to maximise clicks. Cheap, flashy items get clicked far more than expensive, well-made ones.

Click rate can climb every week while average order value quietly falls. The model is doing exactly what it was optimised to do: get clicks, not revenue.

The honest part

There is rarely one perfect metric that captures everything a business cares about. Every metric is a simplification, chosen because it is measurable now.

The goal is not finding a flawless metric. It is watching more than one, and noticing quickly when they start pointing in different directions.

Remember this

  • A model metric and a business metric can move in opposite directions.
  • A model optimised hard for one number can find shortcuts that hurt everything else.
  • Watch both together. A single improving metric is not proof anything real improved.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy pandas

Two rankings of the same items, two different winners

Items where flashy, low-margin products get clicked more than solid, higher-margin ones — a believable, common pattern. Compare picking the top five by predicted click-through against picking the top five by expected revenue.

business_metric.py
import numpy as np
import pandas as pd

rng = np.random.RandomState(0)

n = 200
# Flashy, low-margin items tend to get clicked a lot (click-bait).
# Solid, higher-margin items get clicked less often, but are worth more
# when they do sell. This trade-off is what makes the two metrics disagree.
margin = rng.uniform(50, 500, n)
click_prob = np.clip(0.9 - 0.0015 * margin + rng.normal(0, 0.05, n), 0.02, 0.95)

items = pd.DataFrame({"margin": margin.round(0), "click_prob": click_prob.round(3)})
items["expected_revenue"] = items["click_prob"] * items["margin"]

SHOW = 5

by_ctr = items.sort_values("click_prob", ascending=False).head(SHOW)
by_value = items.sort_values("expected_revenue", ascending=False).head(SHOW)


def summarise(name, chosen):
    mean_ctr = chosen["click_prob"].mean()
    total_expected_revenue = chosen["expected_revenue"].sum()
    print(f"{name:<22} mean predicted click-through: {mean_ctr:.3f}   "
          f"expected revenue for these {SHOW}: {total_expected_revenue:,.0f}")


summarise("Ranked by CTR", by_ctr)
summarise("Ranked by expected revenue", by_value)
Output
Ranked by CTR          mean predicted click-through: 0.865   expected revenue for these 5: 277
Ranked by expected revenue mean predicted click-through: 0.498   expected revenue for these 5: 843

Exact output from this seeded script. Ranking by click-through wins the model metric by a wide margin, and loses the business metric by an even wider one.

Line-by-line walkthrough

click_prob is built to fall as margin rises, with some noise. This one line encodes the entire "clicky but cheap versus solid but valuable" trade-off used throughout the example.

by_ctr picks the five items a click-optimised model would show. by_value picks the five an expected-revenue-optimised model would show. Same catalogue, same size selection, opposite priority.

Nothing here is a bug. Both selections are doing exactly what they were asked to do. The lesson is in which question was asked.

Common mistakes

Optimising a model for the only metric that is easy to measure at prediction time. Clicks happen in milliseconds; a sale might happen minutes later, or not at all. Easy-to-measure and important-to-the-business are not the same property.

Declaring victory when the model metric improves, without checking the business metric. Alerts people do not ignore should watch both, not only the one the model was trained on.

Assuming the two metrics were aligned once, so they still are. The relationship between them can drift on its own, independent of any drift in the model or the data — a slow shift in what users value counts too.

Picking a single business metric and ignoring the ones it can trade against. Revenue can rise while trust falls. Watch a small basket of business metrics together, not one.

Try it yourself

Change the noise in click_prob from rng.normal(0, 0.05, n) to rng.normal(0, 0.15, n). Rerun, and check whether the gap between the two rankings' business-metric outcomes shrinks as the trade-off gets noisier.

What to learn next

Researcher — Mathematics and papers.

This is Goodhart's Law, operationally

"When a measure becomes a target, it ceases to be a good measure" (Goodhart, 1975, in the original economic context; popularised for ML by Strathern, 1997). A model metric optimised directly via gradient descent is a target in the strongest possible sense — the optimiser will exploit any gap between the metric and the true objective with more thoroughness than a human ever would.

Reward hacking and specification gaming

The reinforcement learning literature has documented this pattern extensively under reward hacking: an agent finds a way to score highly on a stated reward function that diverges from the designer's actual intent. Krakovna et al. (2020) maintain a catalogue of concrete specification-gaming examples across RL systems, useful for building intuition for how varied and non-obvious these failure modes are.

Formalising the metric gap

Let $J$ be the true business objective (unmeasurable or measurable only with long delay) and $\hat{J}$ be the proxy metric actually optimised. The correlation $\rho(J, \hat{J})$ under the training distribution can be high while the correlation under the optimised distribution collapses, because optimisation actively seeks out the region where $\hat{J}$ is high and $J$ is not — precisely the region a correlation computed on pre-optimisation data never samples.

$$\hat{J}^* = \arg\max_\pi \, \mathbb{E}{\pi}[\hat{J}] \quad \not\Rightarrow \quad \pi \approx \arg\max\pi \, \mathbb{E}_{\pi}[J]$$

This is why offline correlation between a proxy and a true objective, however strong, is not evidence that optimising the proxy is safe — the correlation is a property of the data the model has seen, not of the model's own optimised policy.

Mitigations used in practice

  • Multi-objective optimisation — optimise a weighted or constrained combination of proxy and (delayed, noisier) business signals, rather than the proxy alone.
  • Counterfactual / long-horizon metrics — measure outcomes over a longer window (session value, weekly retention) that is harder to game with a single-step shortcut.
  • Interleaving and holdouts — releasing a model gradually against a true business metric, not only an offline proxy, before shipping to all traffic.
  • Guardrail metrics — a set of business metrics that must not regress beyond a bound, checked automatically alongside the primary optimisation target.

Papers

  • Goodhart, Problems of Monetary Management: The U.K. Experience, 1975 — the origin of the law.
  • Strathern, 'Improving Ratings': Audit in the British University System, 1997 — the widely cited restatement.
  • Krakovna et al., Specification Gaming: The Flip Side of AI Ingenuity, DeepMind 2020 — deepmind.google/discover/blog/specification-gaming-the-flip-side-of-ai-ingenuity
  • Hadfield-Menell and Hadfield, Incomplete Contracting and AI Alignment, AIES 2019 — frames the proxy-versus-true-objective gap as an incomplete-contract problem.

What to learn next