Observability for LLM Applications

Replaying production traffic

Replaying traffic means running a candidate prompt or model against real, previously logged requests, so a fix gets checked against real questions before it ever reaches a real user.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Replaying traffic means running a candidate fix against real, previously logged requests, before it ever reaches a real user.

The analogy you have already lived

A cricket team that lost a match does not only discuss what should have happened. They sit down and watch the actual recorded footage — the real balls bowled, the real shots played.

Trying a new tactic against imagined, hypothetical situations is a weak substitute. Watching it against what genuinely happened is what actually reveals whether the fix would have worked.

Replaying production traffic gives a prompt fix the same real footage to be tested against, instead of a handful of hand-written guesses.

Why it exists

Earlier in this section, every real request got logged, redacted, and kept. That log is not only a record. It is a ready-made, realistic test set.

A hand-written test set reflects what a developer imagined a user might ask. Real logged traffic reflects what users actually asked, including the strange, unexpected questions nobody thought to write down.

How it works

   real, logged requests (already redacted)
              |
              v
   [ run the CANDIDATE fix on each one ]
              |
              v
   compare: did the candidate do better than what
             actually shipped, on these same real inputs?

Nothing here is invented. Every input the candidate is tested against is something a real user genuinely asked.

A real example you have seen

A weather app testing a new "explain the forecast in plain words" prompt does not only try three sentences a developer typed. It runs the new prompt against a week of real, logged questions. The two versions get compared side by side.

The honest part

Replay tests what a candidate does on the past. It cannot promise the same result on tomorrow's questions, which will not be identical to yesterday's.

Treat a strong replay result as a reason to ship with more confidence. Pair it with live evaluation after launch, not as a reason to stop watching once it is live.

Remember this

  • Replay runs a candidate fix against real, previously logged requests.
  • Real traffic reveals the strange questions a hand-written test set would never think to include.
  • A good replay result is reassuring, not a guarantee about tomorrow's traffic.

What to learn next

Developer — Code and libraries.

Setup

No install needed. This script only uses Python's standard library.

Testing a fix against the exact requests that broke before

Reusing the four real "hour 2" requests from running evaluations on live traffic — the ones the broken prompt version answered badly — replay a candidate fix against those same real inputs.

replay.py
# Real, previously-logged requests -- the exact inputs the broken hour-2
# prompt version received in production. Replaying means running a NEW
# candidate over these same real inputs, never inventing fresh test cases.
logged_requests = [
    {"order_id": "OD1130", "context": "order OD1130, delayed by 1 day"},
    {"order_id": "OD1131", "context": "order OD1131, delivered today"},
    {"order_id": "OD1132", "context": "order OD1132, out for delivery"},
    {"order_id": "OD1133", "context": "order OD1133, refund processed"},
]

# What the OLD (broken) prompt version actually produced for these same
# requests, taken from the production log.
old_responses = [
    "We understand delays can be frustrating and we're always improving our service.",
    "Thank you for your patience with your recent purchase experience.",
    "Your delivery is being handled with care by our logistics team.",
    "OD1133 refund processed. We appreciate your business with us today.",
]


def candidate_response(order_id, context):
    """Stands in for calling the model with the CANDIDATE prompt version.
    A real replay sends `context` to the new prompt and records what the
    model actually says; here a small rule-based rewrite plays that role
    so the example runs without an API key."""
    detail = context.split(", ", 1)[1]
    return f"Update on {order_id}: {detail}."


def check_mentions_order_id(order_id, response):
    return order_id in response


def check_no_generic_filler(response):
    filler = ["always improving", "appreciate your business", "with care by our",
              "patience with your recent"]
    return not any(p in response.lower() for p in filler)


def pass_rate(responses, requests):
    mentions = sum(check_mentions_order_id(r["order_id"], resp)
                   for r, resp in zip(requests, responses))
    no_filler = sum(check_no_generic_filler(resp) for resp in responses)
    n = len(requests)
    return mentions / n, no_filler / n


old_mentions, old_filler = pass_rate(old_responses, logged_requests)
print(f"OLD prompt version -- mentions_order_id: {old_mentions:.0%}   no_generic_filler: {old_filler:.0%}")

new_responses = [candidate_response(r["order_id"], r["context"]) for r in logged_requests]
new_mentions, new_filler = pass_rate(new_responses, logged_requests)
print(f"NEW prompt version -- mentions_order_id: {new_mentions:.0%}   no_generic_filler: {new_filler:.0%}")

print("\nreplayed outputs from the candidate version:")
for r, resp in zip(logged_requests, new_responses):
    print(f"  {r['order_id']}: {resp}")
Output
OLD prompt version -- mentions_order_id: 25%   no_generic_filler: 0%
NEW prompt version -- mentions_order_id: 100%   no_generic_filler: 100%

replayed outputs from the candidate version:
  OD1130: Update on OD1130: delayed by 1 day.
  OD1131: Update on OD1131: delivered today.
  OD1132: Update on OD1132: out for delivery.
  OD1133: Update on OD1133: refund processed.

Exact output from this fixed data and this deterministic candidate function. Both check rates jumped to 100% on the exact same real inputs the old version failed on.

Line-by-line walkthrough

logged_requests and old_responses are treated as historical fact here, standing in for rows pulled from a real, redacted production log.

candidate_response is the one function a real replay harness swaps out. Its body here is a simple, deterministic rewrite; in production it would be a real call using the candidate prompt version.

The pass_rate function is reused unchanged from the online-evals lesson, on purpose. Replay is not a new evaluation technique — it is the same checks, applied to a candidate instead of to what is currently live.

Common mistakes

Replaying against a hand-picked, easy subset of traffic. Cherry-picked inputs make every candidate look good. Replay against a genuinely representative sample, including the traffic that broke things before.

Never redacting the replay set. A replay log is still a log of real user data. Redacting personal data applies to it exactly as it applies to any other log.

Treating a passing replay as equivalent to a passing live rollout. Replay cannot observe how the candidate performs under real production timing, load, or truly novel questions that have not been asked yet. Ship behind a small, watched rollout even after a clean replay.

Comparing old and new on different requests. The entire value of replay comes from holding the input fixed and changing only the candidate. Comparing different samples reintroduces the exact confound replay exists to remove.

Try it yourself

Add a fifth logged request whose context does not follow the "order X, detail" shape the candidate function assumes. Rerun, and see how the candidate handles an input its simple rewrite logic was not built for.

What to learn next

Researcher — Mathematics and papers.

Replay as a form of counterfactual evaluation

Replay answers a narrower, more tractable question than a full counterfactual: "what would the candidate have output, given the same input the production system actually received?" It does not answer "what would the user have done differently, faced with the candidate's output instead?" — that second question requires either a live rollout or the off-policy evaluation techniques discussed in feedback loops in production.

Non-determinism complicates replay for real models

The developer example is deterministic by construction. A real LLM call, even at temperature zero, is not always bit-for-bit reproducible across providers or model versions, due to non-associative floating-point reduction order in batched inference and provider-side infrastructure differences. Replay comparisons against a real model should therefore report a distribution over repeated replays (mean pass rate, variance) rather than treating a single replay run as exact.

Replay sets as a living regression suite

The strongest production pattern treats the replay set as a continuously growing regression suite: every real incident that reveals a failure gets its triggering request added permanently, the way golden examples in CI work for traditional ML. Over time this converges toward a replay set specifically enriched with the historically hardest cases, which is more informative per example than a uniformly random sample of traffic.

Shadow evaluation versus replay

Replay runs the candidate offline against static, already-logged requests. Shadow deployment (covered generally in model serving) runs the candidate online, in real time, against live traffic, without serving its output to the user. Shadow catches issues replay cannot — real latency under real load, real infrastructure interactions — at the cost of running the candidate live, which replay avoids entirely. The two are complementary stages, not substitutes: replay first, on history; shadow second, in real time; then a small live rollout.

Statistical comparison of replay results

When comparing pass rates between old and candidate on the same replayed set, a paired test (McNemar's test, for binary pass/fail outcomes on matched pairs) is the correct choice over an unpaired two-proportion test, because both versions are evaluated on the identical set of inputs rather than independent samples — pairing removes item-to-item difficulty variance from the comparison, substantially increasing statistical power at a fixed sample size.

Papers

  • McNemar, Note on the Sampling Error of the Difference Between Correlated Proportions or Percentages, Psychometrika 1947 — the paired test referenced above.
  • Ribeiro et al., Beyond Accuracy: Behavioral Testing of NLP Models with CheckList, ACL 2020 — the regression-suite-that-grows-from-real-failures pattern, applied generally to NLP systems.

What to learn next