Observability for LLM Applications
Replaying production traffic
Replaying traffic means running a candidate prompt or model against real, previously logged requests, so a fix gets checked against real questions before it ever reaches a real user.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Replaying traffic means running a candidate fix against real, previously logged requests, before it ever reaches a real user.
The analogy you have already lived
A cricket team that lost a match does not only discuss what should have happened. They sit down and watch the actual recorded footage — the real balls bowled, the real shots played.
Trying a new tactic against imagined, hypothetical situations is a weak substitute. Watching it against what genuinely happened is what actually reveals whether the fix would have worked.
Replaying production traffic gives a prompt fix the same real footage to be tested against, instead of a handful of hand-written guesses.
Why it exists
Earlier in this section, every real request got logged, redacted, and kept. That log is not only a record. It is a ready-made, realistic test set.
A hand-written test set reflects what a developer imagined a user might ask. Real logged traffic reflects what users actually asked, including the strange, unexpected questions nobody thought to write down.
How it works
real, logged requests (already redacted)
|
v
[ run the CANDIDATE fix on each one ]
|
v
compare: did the candidate do better than what
actually shipped, on these same real inputs?Nothing here is invented. Every input the candidate is tested against is something a real user genuinely asked.
A real example you have seen
A weather app testing a new "explain the forecast in plain words" prompt does not only try three sentences a developer typed. It runs the new prompt against a week of real, logged questions. The two versions get compared side by side.
The honest part
Replay tests what a candidate does on the past. It cannot promise the same result on tomorrow's questions, which will not be identical to yesterday's.
Treat a strong replay result as a reason to ship with more confidence. Pair it with live evaluation after launch, not as a reason to stop watching once it is live.
Remember this
- Replay runs a candidate fix against real, previously logged requests.
- Real traffic reveals the strange questions a hand-written test set would never think to include.
- A good replay result is reassuring, not a guarantee about tomorrow's traffic.
What to learn next
- Observability for agent runs — replaying gets harder once a request can take a different path through tools each time.
- Running evaluations on live traffic — the same checks used here, run continuously after a candidate ships.
- CI/CD for machine learning — building a replay set into an automated pipeline, the way a golden test set works for a traditional model.
Developer — Code and libraries.
Setup
No install needed. This script only uses Python's standard library.
Testing a fix against the exact requests that broke before
Reusing the four real "hour 2" requests from running evaluations on live traffic — the ones the broken prompt version answered badly — replay a candidate fix against those same real inputs.
# Real, previously-logged requests -- the exact inputs the broken hour-2
# prompt version received in production. Replaying means running a NEW
# candidate over these same real inputs, never inventing fresh test cases.
logged_requests = [
{"order_id": "OD1130", "context": "order OD1130, delayed by 1 day"},
{"order_id": "OD1131", "context": "order OD1131, delivered today"},
{"order_id": "OD1132", "context": "order OD1132, out for delivery"},
{"order_id": "OD1133", "context": "order OD1133, refund processed"},
]
# What the OLD (broken) prompt version actually produced for these same
# requests, taken from the production log.
old_responses = [
"We understand delays can be frustrating and we're always improving our service.",
"Thank you for your patience with your recent purchase experience.",
"Your delivery is being handled with care by our logistics team.",
"OD1133 refund processed. We appreciate your business with us today.",
]
def candidate_response(order_id, context):
"""Stands in for calling the model with the CANDIDATE prompt version.
A real replay sends `context` to the new prompt and records what the
model actually says; here a small rule-based rewrite plays that role
so the example runs without an API key."""
detail = context.split(", ", 1)[1]
return f"Update on {order_id}: {detail}."
def check_mentions_order_id(order_id, response):
return order_id in response
def check_no_generic_filler(response):
filler = ["always improving", "appreciate your business", "with care by our",
"patience with your recent"]
return not any(p in response.lower() for p in filler)
def pass_rate(responses, requests):
mentions = sum(check_mentions_order_id(r["order_id"], resp)
for r, resp in zip(requests, responses))
no_filler = sum(check_no_generic_filler(resp) for resp in responses)
n = len(requests)
return mentions / n, no_filler / n
old_mentions, old_filler = pass_rate(old_responses, logged_requests)
print(f"OLD prompt version -- mentions_order_id: {old_mentions:.0%} no_generic_filler: {old_filler:.0%}")
new_responses = [candidate_response(r["order_id"], r["context"]) for r in logged_requests]
new_mentions, new_filler = pass_rate(new_responses, logged_requests)
print(f"NEW prompt version -- mentions_order_id: {new_mentions:.0%} no_generic_filler: {new_filler:.0%}")
print("\nreplayed outputs from the candidate version:")
for r, resp in zip(logged_requests, new_responses):
print(f" {r['order_id']}: {resp}")OLD prompt version -- mentions_order_id: 25% no_generic_filler: 0% NEW prompt version -- mentions_order_id: 100% no_generic_filler: 100% replayed outputs from the candidate version: OD1130: Update on OD1130: delayed by 1 day. OD1131: Update on OD1131: delivered today. OD1132: Update on OD1132: out for delivery. OD1133: Update on OD1133: refund processed.
Exact output from this fixed data and this deterministic candidate function. Both check rates jumped to 100% on the exact same real inputs the old version failed on.
Line-by-line walkthrough
logged_requests and old_responses are treated as historical fact here, standing in for rows pulled from a real, redacted production log.
candidate_response is the one function a real replay harness swaps out. Its body here is a simple, deterministic rewrite; in production it would be a real call using the candidate prompt version.
The pass_rate function is reused unchanged from the online-evals lesson, on purpose. Replay is not a new evaluation technique — it is the same checks, applied to a candidate instead of to what is currently live.
Common mistakes
Replaying against a hand-picked, easy subset of traffic. Cherry-picked inputs make every candidate look good. Replay against a genuinely representative sample, including the traffic that broke things before.
Never redacting the replay set. A replay log is still a log of real user data. Redacting personal data applies to it exactly as it applies to any other log.
Treating a passing replay as equivalent to a passing live rollout. Replay cannot observe how the candidate performs under real production timing, load, or truly novel questions that have not been asked yet. Ship behind a small, watched rollout even after a clean replay.
Comparing old and new on different requests. The entire value of replay comes from holding the input fixed and changing only the candidate. Comparing different samples reintroduces the exact confound replay exists to remove.
Try it yourself
Add a fifth logged request whose context does not follow the "order X, detail" shape the candidate function assumes. Rerun, and see how the candidate handles an input its simple rewrite logic was not built for.
What to learn next
- Observability for agent runs — replaying gets harder once a request can take a different path through tools each time.
- Running evaluations on live traffic — the same checks used here, run continuously after a candidate ships.
- CI/CD for machine learning — building a replay set into an automated pipeline, the way a golden test set works for a traditional model.
Researcher — Mathematics and papers.
Replay as a form of counterfactual evaluation
Replay answers a narrower, more tractable question than a full counterfactual: "what would the candidate have output, given the same input the production system actually received?" It does not answer "what would the user have done differently, faced with the candidate's output instead?" — that second question requires either a live rollout or the off-policy evaluation techniques discussed in feedback loops in production.
Non-determinism complicates replay for real models
The developer example is deterministic by construction. A real LLM call, even at temperature zero, is not always bit-for-bit reproducible across providers or model versions, due to non-associative floating-point reduction order in batched inference and provider-side infrastructure differences. Replay comparisons against a real model should therefore report a distribution over repeated replays (mean pass rate, variance) rather than treating a single replay run as exact.
Replay sets as a living regression suite
The strongest production pattern treats the replay set as a continuously growing regression suite: every real incident that reveals a failure gets its triggering request added permanently, the way golden examples in CI work for traditional ML. Over time this converges toward a replay set specifically enriched with the historically hardest cases, which is more informative per example than a uniformly random sample of traffic.
Shadow evaluation versus replay
Replay runs the candidate offline against static, already-logged requests. Shadow deployment (covered generally in model serving) runs the candidate online, in real time, against live traffic, without serving its output to the user. Shadow catches issues replay cannot — real latency under real load, real infrastructure interactions — at the cost of running the candidate live, which replay avoids entirely. The two are complementary stages, not substitutes: replay first, on history; shadow second, in real time; then a small live rollout.
Statistical comparison of replay results
When comparing pass rates between old and candidate on the same replayed set, a paired test (McNemar's test, for binary pass/fail outcomes on matched pairs) is the correct choice over an unpaired two-proportion test, because both versions are evaluated on the identical set of inputs rather than independent samples — pairing removes item-to-item difficulty variance from the comparison, substantially increasing statistical power at a fixed sample size.
Papers
- McNemar, Note on the Sampling Error of the Difference Between Correlated Proportions or Percentages, Psychometrika 1947 — the paired test referenced above.
- Ribeiro et al., Beyond Accuracy: Behavioral Testing of NLP Models with CheckList, ACL 2020 — the regression-suite-that-grows-from-real-failures pattern, applied generally to NLP systems.
What to learn next
- Observability for agent runs — replaying gets harder once a request can take a different path through tools each time.
- Running evaluations on live traffic — the same checks used here, run continuously after a candidate ships.
- CI/CD for machine learning — building a replay set into an automated pipeline, the way a golden test set works for a traditional model.