Observability for LLM Applications
Running evaluations on live traffic
An online eval is an automatic check run on real answers as they happen, catching a quality drop within hours instead of whenever someone happens to notice.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
An online eval is an automatic check run on real answers as they happen, not only on a test set before launch.
An eval here is short for evaluation — a check that scores how good an answer is, against some rule or standard.
The analogy you have already lived
A learner driver passes a test on a closed track, once, before getting a licence. That single pass says nothing about how they drive on a real, busy road every day after.
A good driving instructor rides along on real trips too, watching real driving. A bad habit gets caught within days, not at the next renewal, years later.
Testing an LLM app once before launch is the closed track. Running checks on its real, live answers is the instructor riding along.
Why it exists
A prompt or model change gets tested against a fixed set of examples before shipping. That test set is small, and it is frozen — it cannot represent every real question a user will actually ask.
Real traffic is much larger and much stranger than any test set. A change that passed every offline test can still behave badly on the odd real questions nobody thought to write down.
How it works
real answer goes out to the user
|
v
[ automatic checks run on it, right away ]
- is it a reasonable length?
- does it actually mention what it should?
- does it dodge the question with stock phrases?
|
v
pass rate tracked over time, hour by hourNone of these checks need a human to read every answer. They run instantly, on every single response.
A real example you have seen
A support bot that always used to name the order number starts giving vague, generic responses after a small prompt tweak. An automatic check for "did it mention the order number?" catches that within the hour. Nobody had to wait for angry customers to complain.
The honest part
Automatic checks like these are rules, not understanding. They can be fooled, and they can flag something that is actually fine.
Treat them as a fast, cheap first filter. The next lesson covers a check with more judgment, for when a simple rule is not enough.
Remember this
- An online eval checks real, live answers automatically, as they happen.
- Offline tests before launch cannot represent everything real traffic will ask.
- Simple automatic checks are a fast first filter, not the final word on quality.
What to learn next
- LLM-as-judge in production — a check with real judgment, for what simple rules cannot catch.
- Capturing user feedback — the signal that comes directly from the people receiving these answers.
- Alerts people do not ignore — turning a falling pass rate like hour 2's into something that pages a human.
Developer — Code and libraries.
Setup
No install needed. This script only uses Python's standard library.
Three simple checks, run on eight live-looking responses
Two batches of support-bot replies, the second one after a prompt change. Three automatic checks, none of them needing a model or a human.
# Simulated live responses from a support bot. Each has the context it was
# given (with a real order number) and what the model actually replied.
# hour_2 is deliberately worse: a prompt change that morning made answers
# ramble and stop mentioning the order number.
live_traffic = [
# hour 1 -- normal
{"hour": 1, "order_id": "OD1123", "context": "order OD1123, delayed by 2 days",
"response": "Your order OD1123 is delayed by 2 days. We're sorry for the wait."},
{"hour": 1, "order_id": "OD1124", "context": "order OD1124, delivered yesterday",
"response": "Order OD1124 was delivered yesterday. Let us know if it didn't arrive."},
{"hour": 1, "order_id": "OD1125", "context": "order OD1125, out for delivery",
"response": "OD1125 is out for delivery today."},
{"hour": 1, "order_id": "OD1126", "context": "order OD1126, refund processed",
"response": "Your refund for OD1126 has been processed."},
# hour 2 -- after the prompt change
{"hour": 2, "order_id": "OD1130", "context": "order OD1130, delayed by 1 day",
"response": "We understand delays can be frustrating and we're always improving our service."},
{"hour": 2, "order_id": "OD1131", "context": "order OD1131, delivered today",
"response": "Thank you for your patience with your recent purchase experience."},
{"hour": 2, "order_id": "OD1132", "context": "order OD1132, out for delivery",
"response": "Your delivery is being handled with care by our logistics team."},
{"hour": 2, "order_id": "OD1133", "context": "order OD1133, refund processed",
"response": "OD1133 refund processed. We appreciate your business with us today."},
]
def check_mentions_order_id(row):
"""Grounding check: did the reply actually reference the real order?"""
return row["order_id"] in row["response"]
def check_reasonable_length(row):
"""Sanity check: not empty, not a wall of text."""
words = len(row["response"].split())
return 3 <= words <= 30
def check_no_generic_filler(row):
"""Heuristic: flags stock corporate phrases that dodge the actual question."""
filler = ["we're always improving", "we appreciate your business", "with care by our"]
return not any(phrase in row["response"].lower() for phrase in filler)
CHECKS = {
"mentions_order_id": check_mentions_order_id,
"reasonable_length": check_reasonable_length,
"no_generic_filler": check_no_generic_filler,
}
for hour in (1, 2):
rows = [r for r in live_traffic if r["hour"] == hour]
print(f"hour {hour} ({len(rows)} responses):")
for check_name, check_fn in CHECKS.items():
pass_rate = sum(check_fn(r) for r in rows) / len(rows)
print(f" {check_name:<20} pass rate: {pass_rate:.0%}")
print()hour 1 (4 responses): mentions_order_id pass rate: 100% reasonable_length pass rate: 100% no_generic_filler pass rate: 100% hour 2 (4 responses): mentions_order_id pass rate: 25% reasonable_length pass rate: 100% no_generic_filler pass rate: 25%
Every number here is exact, since every check is a deterministic rule applied to fixed, hand-written text. A real system's traffic and pass rates will differ, and will genuinely vary from what any two runs show, because real user questions vary.
Line-by-line walkthrough
check_mentions_order_id is a grounding check — it does not ask whether the reply sounds good, only whether it actually used the real fact it was given.
check_reasonable_length catches the crudest failures: an empty reply, or a wall of text nobody asked for. Notice it stayed at 100% in hour 2 — length alone said nothing was wrong.
check_no_generic_filler is a hand-written list of stock phrases. It is deliberately narrow, and it is exactly the kind of check worth adding the moment a real failure pattern is spotted.
Common mistakes
Relying on one check to catch every kind of problem. reasonable_length passed at 100% in both hours, while two other checks caught a real regression. Run several narrow checks, not one broad one.
Writing checks only after a problem, never before. The value compounds: each real incident should leave behind one more permanent, automatic check, so the same failure cannot slip through silently again.
Treating a 100% pass rate as proof of quality. It only proves the specific things being checked were fine. It says nothing about whatever nobody thought to check for yet.
Running checks only occasionally, on a sample, instead of continuously. A prompt regression that lasts three hours before a scheduled check runs is three hours of bad answers nobody caught in real time.
Try it yourself
Add a fourth check, check_no_question_dodging, that fails if the response does not contain a digit anywhere (a rough proxy for "actually referenced a concrete detail"). Rerun, and compare its pass rate against mentions_order_id.
What to learn next
- LLM-as-judge in production — a check with real judgment, for what simple rules cannot catch.
- Capturing user feedback — the signal that comes directly from the people receiving these answers.
- Alerts people do not ignore — turning a falling pass rate like hour 2's into something that pages a human.
Researcher — Mathematics and papers.
Where rule-based online evals sit in the evaluation stack
Automatic, rule-based checks like the ones above are the cheapest tier of a broader evaluation stack, typically layered as:
- Deterministic checks — regex, schema validation, keyword presence, length bounds. Zero cost beyond compute, zero ambiguity, and blind to anything not explicitly encoded.
- Model-based checks (LLM-as-judge) — covered in the next lesson. Catches nuance the rules above cannot, at real per-call cost and its own failure modes.
- Human review — the ground truth, applied to a sample, too expensive to run on all traffic.
A mature online evaluation system runs tier 1 on 100% of traffic, tier 2 on a sample, and tier 3 on an even smaller, targeted sample (often the cases tiers 1 and 2 flagged as uncertain).
Grounding checks and hallucination proxies
check_mentions_order_id is a minimal instance of a grounding check: verifying that specific factual content from the retrieved context appears in the output. At scale, this generalises to checking that named entities, numbers, and dates in a response are traceable to the provided context — a cheap, high-precision proxy for one specific kind of hallucination (fabricating or dropping a concrete fact), though it says nothing about subtler failures like an incorrect inference drawn correctly-worded from real facts.
Statistical considerations for online metrics
A pass rate computed on a small live sample, as in the developer example (n=4 per hour), carries substantial sampling noise. A Wilson score interval, rather than a naive $\hat{p} \pm 1.96\sqrt{\hat{p}(1-\hat{p})/n}$ normal approximation, is the standard correction for proportions estimated from small $n$, since the normal approximation misbehaves badly near 0% or 100% — exactly the range online eval pass rates often sit in for a healthy system, making the correction more than a theoretical nicety here.
Papers and tools
- Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, NeurIPS 2023 — arxiv.org/abs/2306.05685, positions rule-based and model-based evaluation as complementary tiers.
- Wilson, Probable Inference, the Law of Succession, and Statistical Inference, JASA 1927 — the origin of the Wilson score interval used for small-sample pass rates.
- Ribeiro et al., Beyond Accuracy: Behavioral Testing of NLP Models with CheckList, ACL 2020 — the same "many narrow checks beat one broad metric" principle, applied to NLP model testing generally.
What to learn next
- LLM-as-judge in production — a check with real judgment, for what simple rules cannot catch.
- Capturing user feedback — the signal that comes directly from the people receiving these answers.
- Alerts people do not ignore — turning a falling pass rate like hour 2's into something that pages a human.