Monitoring Models in Production
Feedback loops in production
A feedback loop is when a model's own past predictions shape the data it gets trained on next, so an early lucky guess can lock itself in as "the truth" forever.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A feedback loop is when a model's own past predictions shape the data it later trains on. An early guess can lock itself in as "true".
The analogy you have already lived
A teacher keeps calling on the same three confident students who raise their hands fastest. Those three get more practice answering out loud, and keep getting picked.
The quiet students in the back never get a turn. Not because they know less. They were never sampled, so nobody ever finds out what they know.
Ask that teacher who the "best" students are after a year. The answer says more about who got called on than who actually knew the most.
Why it exists
A recommender only learns from what it showed. If it never showed an item, it never sees anyone react to that item, positively or negatively.
An early, lucky guess about which item is good gets more exposure. More exposure means more clicks, purely from being seen more. More clicks looks like proof the guess was right. The system trains itself into agreeing with its own first impression.
How it works
model recommends item X
|
v
only item X gets shown, so only item X gets feedback
|
v
item X's numbers look great (it's the only one still being tried)
|
v
model recommends item X again, more confidentlyNothing in this loop needs the model to be wrong on purpose. Confirming its own first guess is enough.
A real example you have seen
A video platform's "trending" list can become mostly self-fulfilling. Videos placed there get more views purely from placement. More views look like proof they deserve it. The same videos stay trending.
A genuinely better video that never got an early push may never get the chance to prove itself.
The honest part
This is one of the harder problems in this whole section, because nothing looks broken from inside the loop. Every individual number the model reports can be completely accurate.
The bias is not in any single measurement. It is in which things got measured at all.
Remember this
- A feedback loop forms when a model's own output shapes its future training data.
- An early lucky guess can look "proven correct" purely from getting shown more.
- Nothing looks wrong from inside the loop — every reported number can be accurate.
What to learn next
- When model metrics and business metrics disagree — a feedback loop can make a model metric look great while the real outcome quietly suffers.
- Monitoring by segment — a feedback loop often locks in for one segment before anyone notices.
- Observability for agent runs — the same self-reinforcing pattern shows up when an agent's own tool choices shape what it learns to trust.
Developer — Code and libraries.
Setup
pip install numpyWatching a loop lock in the wrong item
Five items with fixed, true quality scores. A recommender that always shows whichever item currently looks best, and never spends a single impression re-checking the others.
import numpy as np
rng = np.random.RandomState(0)
true_quality = np.array([0.50, 0.52, 0.48, 0.55, 0.51]) # item 3 is genuinely best
n_items = len(true_quality)
BATCH = 20
ROUNDS = 40
# Round 0: every item gets one small trial batch, same as launch day.
shown_count = np.ones(n_items) * BATCH
clicks = rng.binomial(BATCH, true_quality)
observed_rate = clicks / shown_count
exposure = shown_count.copy()
for _ in range(ROUNDS):
# GREEDY recommender: always show whichever item currently LOOKS best.
# It never spends a single impression checking on the others again.
best = int(np.argmax(observed_rate))
new_clicks = rng.binomial(BATCH, true_quality[best])
clicks[best] += new_clicks
shown_count[best] += BATCH
observed_rate[best] = clicks[best] / shown_count[best]
exposure[best] += BATCH
print("item true quality times recommended final observed rate")
for i in range(n_items):
marker = " <- truly best" if i == np.argmax(true_quality) else ""
print(f"item {i} {true_quality[i]:.2f} {int(exposure[i]):>5}"
f" {observed_rate[i]:.3f}{marker}")
winner = int(np.argmax(exposure))
truly_best = int(np.argmax(true_quality))
print(f"\nitem the system converged on: item {winner}")
print(f"item that was actually best: item {truly_best}")item true quality times recommended final observed rate item 0 0.50 400 0.512 item 1 0.52 20 0.450 item 2 0.48 260 0.469 item 3 0.55 60 0.467 <- truly best item 4 0.51 160 0.463 item the system converged on: item 0 item that was actually best: item 3
Exact output from this seeded script. The system locked onto item 0, a middling item, and showed it 400 times. The genuinely best item, item 3, was shown only 60 times and never got the chance to prove itself.
Line-by-line walkthrough
true_quality never changes during the run. It is the ground truth the system is trying, and failing, to discover.
The round-0 block gives every item one honest, equal-sized trial. This is the only fair look any item ever gets in this script.
Inside the loop, best = int(np.argmax(observed_rate)) is the entire bug. The system commits fully to whichever item's early numbers happened to look good, and stops sampling everything else.
Common mistakes
Reading a stable, high metric as proof of a good recommendation. Item 0's 0.512 observed rate looks solid. It is solid, and it is still the wrong answer, because nothing else was ever compared against it fairly again.
Assuming more traffic fixes this on its own. Running the loop above for 400 rounds instead of 40 makes the lock-in worse, not better. Volume reinforces the loop; it does not correct it.
Never holding out a small slice of random exposure. Reserving even 5% of impressions for a random or rotating item continuously re-checks the system's own assumptions, at a small, bounded cost.
Treating this as a modelling problem instead of a data-collection problem. No amount of retraining fixes a feedback loop, because the training data itself is the thing that is biased. The fix is in how exposure is allocated, not in the model.
Try it yourself
Add a line so that 10% of each round's batch goes to a random item instead of the current best, before the rest goes to the greedy choice. Rerun, and check whether item 3 gets discovered.
What to learn next
- When model metrics and business metrics disagree — a feedback loop can make a model metric look great while the real outcome quietly suffers.
- Monitoring by segment — a feedback loop often locks in for one segment before anyone notices.
- Observability for agent runs — the same self-reinforcing pattern shows up when an agent's own tool choices shape what it learns to trust.
Researcher — Mathematics and papers.
This is the exploration-exploitation trade-off
The greedy policy in the developer example is the textbook failure mode of the multi-armed bandit problem: pure exploitation, with zero exploration, converges to whichever arm's early samples were luckiest, with probability that does not go to zero as the number of rounds grows. This is a well-known result, not a coincidence of the seed chosen for this lesson.
Standard corrections
- $\epsilon$-greedy — explore a random arm with fixed probability $\epsilon$, exploit the current best otherwise. Simple, and reserves a fixed, bounded cost for continued discovery.
- Upper Confidence Bound (UCB) — select the arm maximising an optimism-adjusted estimate, $\hat{\mu}_i + c\sqrt{\frac{\ln t}{n_i}}$, where $\hat{\mu}_i$ is the observed rate for arm $i$, $n_i$ its sample count, and $t$ the total round count. Under-sampled arms get a wide confidence bonus, which shrinks as $n_i$ grows.
- Thompson Sampling — maintain a posterior distribution over each arm's true rate (a Beta distribution, for binary outcomes), and sample from it to choose an arm, naturally exploring more where uncertainty is higher.
Auer, Cesa-Bianchi and Fischer (2002) give the canonical regret bounds for $\epsilon$-greedy and UCB, showing UCB achieves logarithmic regret while a purely greedy policy can incur regret linear in the number of rounds, exactly the failure demonstrated above.
The recommender-systems literature on this exact loop
Chaney, Stewart and Engelhardt (2018), How Algorithmic Confounding in Recommendation Systems Increases Homogeneity and Decreases Utility, formalises this as algorithmic confounding: because user behaviour is itself shaped by algorithmic exposure, naively treating logged interactions as i.i.d. samples of "true preference" is a causal-inference error, not only a statistics-hygiene issue.
Counterfactual and off-policy correction
When the logging policy's exposure probabilities are recorded, inverse propensity scoring can reweight logged data to approximate what an unbiased policy would have observed:
$$\hat{V}(\pi) = \frac{1}{N}\sum_{i=1}^{N} \frac{\pi(a_i \mid x_i)}{\pi_0(a_i \mid x_i)}\, r_i$$
Where $\pi_0$ is the logging (production) policy's probability of the action actually taken, $\pi$ is the policy being evaluated, and $r_i$ is the observed reward. This only works if $\pi_0(a_i \mid x_i) > 0$ for every action the evaluated policy might take — a policy with zero exploration, as in the developer example, makes this correction mathematically impossible after the fact. Logging propensities is a decision that has to be made before the loop runs, not recovered afterward.
Papers
- Auer, Cesa-Bianchi and Fischer, Finite-time Analysis of the Multiarmed Bandit Problem, Machine Learning 2002.
- Chapelle and Li, An Empirical Evaluation of Thompson Sampling, NeurIPS 2011.
- Chaney, Stewart and Engelhardt, How Algorithmic Confounding in Recommendation Systems Increases Homogeneity and Decreases Utility, RecSys 2018 — arxiv.org/abs/1710.11214
- Dudík, Langford and Li, Doubly Robust Policy Evaluation and Learning, ICML 2011 — arxiv.org/abs/1103.4601
What to learn next
- When model metrics and business metrics disagree — a feedback loop can make a model metric look great while the real outcome quietly suffers.
- Monitoring by segment — a feedback loop often locks in for one segment before anyone notices.
- Observability for agent runs — the same self-reinforcing pattern shows up when an agent's own tool choices shape what it learns to trust.