Observability for LLM Applications
LLM-as-judge in production
LLM-as-judge means asking a second model to grade the first model's answers against a rubric, catching nuanced problems that a simple rule can never write down.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
LLM-as-judge means asking a second model to grade the first model's answers, using a rubric, instead of a hand-written rule.
The analogy you have already lived
A tailoring shop has junior tailors stitching all day. The owner cannot personally measure every seam on every garment.
Instead, a senior tailor with a trained eye checks a sample of finished pieces. They judge fit and finish the way a measuring tape never could, by knowing what "good" looks like.
Using one model to grade another model's answers is that senior tailor's eye, applied to text instead of cloth.
Why it exists
The previous lesson covered simple, rule-based checks. They are fast and cheap, and they are also narrow — a rule can only catch what someone thought to write a rule for.
Some questions need judgment. Is this answer actually helpful, or technically correct but useless? Is the tone right? Did it politely dodge the real question? A rule struggles here. A model reading the answer, against clear instructions, does much better.
How it works
real answer ----> [ judge model reads it, plus a rubric ] ----> a score + reasoning
"was this accurate?
was this helpful?
score 1-5, explain why"The judge never sees the question alone. It sees the context, the answer, and clear grading instructions, the same way a human grader would.
A real example you have seen
A shopping assistant's answers can all pass simple checks: right length, no banned phrases. Some are still subtly unhelpful, technically correct but missing the point. A judge model, asked "did this actually help?", catches that gap.
The honest part
A judge model is still a model. It can be fooled, it can be inconsistent between runs, and it can share the same blind spots as the model it is grading.
Treat its score as a strong, useful signal — not as ground truth. Capturing user feedback from real people is what eventually calibrates whether the judge itself is any good.
Remember this
- LLM-as-judge uses a second model, plus a rubric, to score answers a rule cannot.
- It catches nuance — helpfulness, tone, subtle dodging — that keyword checks miss.
- A judge's score is a strong signal, not the final truth.
What to learn next
- Capturing user feedback — the human signal a judge model should eventually be checked against.
- Token and cost telemetry — accounting for the extra cost a judge call adds.
- Running evaluations on live traffic — the cheaper, rule-based layer a judge model sits on top of.
Developer — Code and libraries.
Setup
No install needed for this example. A real judge call needs whichever provider's SDK you use, plus an API key.
Seeing the shape of it, without an API key
Grading needs a real model call in production. To keep this example runnable for everyone, mock_judge below stands in for that call — its body is what a real system replaces, not the calling code around it. The real prompt it would send is shown too.
# The prompt a real LLM-as-judge call would send. In production this string
# is sent to a model API; here it is shown but never sent anywhere.
JUDGE_PROMPT = """You are grading a customer support reply.
Context given to the assistant: {context}
Assistant's reply: {response}
Score the reply from 1 (bad) to 5 (excellent) on:
- accuracy: does it use the real facts from the context?
- helpfulness: does it actually answer, without vague filler?
Reply with JSON: {{"score": <1-5>, "reasoning": "<one sentence>"}}"""
def mock_judge(context, response):
"""
Stands in for a real call to an LLM judge, so this example runs with
no API key. A real system replaces this function's BODY with an API
call using JUDGE_PROMPT above -- the calling code does not change.
"""
mentions_fact = any(tok in response for tok in context.split() if tok.isupper() or tok.isdigit())
filler_phrases = ["always improving", "appreciate your business", "with care by our",
"patience with your recent"]
has_filler = any(p in response.lower() for p in filler_phrases)
word_count = len(response.split())
score = 5
reasons = []
if not mentions_fact:
score -= 2
reasons.append("does not reference a concrete fact from the context")
if has_filler:
score -= 2
reasons.append("leans on generic filler instead of the specifics")
if word_count > 30:
score -= 1
reasons.append("longer than it needs to be")
score = max(1, score)
reasoning = "; ".join(reasons) if reasons else "grounded, concise, and on-topic"
return {"score": score, "reasoning": reasoning}
examples = [
{"context": "order OD1123, delayed by 2 days",
"response": "Your order OD1123 is delayed by 2 days. We're sorry for the wait."},
{"context": "order OD1130, delayed by 1 day",
"response": "We understand delays can be frustrating and we're always improving our service."},
{"context": "order OD1131, delivered today",
"response": "Thank you for your patience with your recent purchase experience."},
]
for ex in examples:
verdict = mock_judge(ex["context"], ex["response"])
print(f"response: {ex['response']!r}")
print(f" -> score: {verdict['score']}/5 reasoning: {verdict['reasoning']}")
print()
print("prompt a real judge model would receive for the first example:\n")
print(JUDGE_PROMPT.format(context=examples[0]["context"], response=examples[0]["response"]))response: "Your order OD1123 is delayed by 2 days. We're sorry for the wait."
-> score: 5/5 reasoning: grounded, concise, and on-topic
response: "We understand delays can be frustrating and we're always improving our service."
-> score: 1/5 reasoning: does not reference a concrete fact from the context; leans on generic filler instead of the specifics
response: 'Thank you for your patience with your recent purchase experience.'
-> score: 1/5 reasoning: does not reference a concrete fact from the context; leans on generic filler instead of the specifics
prompt a real judge model would receive for the first example:
You are grading a customer support reply.
Context given to the assistant: order OD1123, delayed by 2 days
Assistant's reply: Your order OD1123 is delayed by 2 days. We're sorry for the wait.
Score the reply from 1 (bad) to 5 (excellent) on:
- accuracy: does it use the real facts from the context?
- helpfulness: does it actually answer, without vague filler?
Reply with JSON: {"score": <1-5>, "reasoning": "<one sentence>"}mock_judge's exact scores are deterministic, since it is plain rule-based code, not a real model call. A real judge model's scores would vary slightly between calls, even on identical input — that variability is a genuine property of using a real judge, not something to fight, only to measure and account for.
Line-by-line walkthrough
JUDGE_PROMPT is written the way a rubric for a human grader would be written: clear criteria, a fixed scale, a requested output format. A real judge call sends this, filled in, to a model, and parses the JSON it returns.
mock_judge only exists so this page has real, runnable output. Its heuristics are deliberately similar to the previous lesson's rule-based checks, to make an honest point: a judge model can be thought of as automating and generalising exactly this kind of reasoning, at the cost of an extra model call.
Common mistakes
Using the same model to judge itself. A model can share its own blind spots with itself, and researchers have found models rate their own outputs more favourably. Use a different, typically stronger, model as judge where possible.
Never checking the judge against human ratings. A judge model that has not been checked against real human judgment on a sample is an unvalidated assumption, not a trusted metric.
Asking for one overall score instead of a rubric per criterion. "Rate this 1-5" alone is vague even to a human. Separate accuracy, helpfulness, and tone into their own scored criteria, the way JUDGE_PROMPT does.
Ignoring cost. Every judged response is a second model call. Judging 100% of traffic can double your total spend, covered in token and cost telemetry. Judging a sample is usually the right trade-off.
Try it yourself
Add a too_short reason to mock_judge that fires when word_count < 3, and lower the score for it. Test it against an empty-ish response like "".
What to learn next
- Capturing user feedback — the human signal a judge model should eventually be checked against.
- Token and cost telemetry — accounting for the extra cost a judge call adds.
- Running evaluations on live traffic — the cheaper, rule-based layer a judge model sits on top of.
Researcher — Mathematics and papers.
Known biases in LLM-as-judge
Zheng et al. (2023), Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, is the foundational study, and documents several systematic biases worth designing around:
- Position bias — when shown two answers to compare, judges favour whichever is presented first, at rates well above chance. Mitigated by evaluating both orderings and averaging, or randomising order per call.
- Verbosity bias — longer answers are rated as better independent of actual quality. Mitigated by explicitly instructing the judge to penalise unnecessary length, and by spot-checking with human raters.
- Self-enhancement bias — a model tends to rate outputs from its own family more favourably. Mitigated by using a judge from a different model family than the one being evaluated.
Correlating a judge with human judgment
Before trusting a judge in production, measure its agreement with human ratings on a labelled sample, using Cohen's kappa for categorical scores or Spearman's rank correlation for ordinal ones, rather than raw agreement percentage, which does not correct for chance agreement:
$$\kappa = \frac{p_o - p_e}{1 - p_e}$$
Where $p_o$ is the observed agreement rate between judge and human, and $p_e$ is the agreement rate expected purely by chance given each rater's marginal distribution of scores. Zheng et al. report GPT-4-as-judge agreement with human preference in the 80-85% range on their benchmark — a useful anchor figure, not a guarantee that transfers to a different domain or rubric without its own validation.
Pairwise versus pointwise judging
Pointwise scoring (as in the developer example: one response, one absolute score) is cheaper but less reliable, since absolute scales are harder for any grader, human or model, to apply consistently across sessions. Pairwise comparison ("which of these two responses is better?") is more reliable per comparison and is the basis of Chatbot Arena-style Elo ranking, at the cost of needing $O(n^2)$ comparisons to fully rank $n$ candidates rather than $O(n)$ scores.
Calibrating and correcting judge scores
Because judge scores carry both systematic bias and random noise, some production systems apply a correction model: fit a small regression from judge score (plus known bias-correlated features like response length) to human score, on a labelled calibration set, and apply that correction to judge scores at scale rather than trusting them raw.
Papers
- Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, NeurIPS 2023 — arxiv.org/abs/2306.05685
- Zheng et al., Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference, ICML 2024
- Cohen, A Coefficient of Agreement for Nominal Scales, Educational and Psychological Measurement 1960 — the origin of Cohen's kappa.
What to learn next
- Capturing user feedback — the human signal a judge model should eventually be checked against.
- Token and cost telemetry — accounting for the extra cost a judge call adds.
- Running evaluations on live traffic — the cheaper, rule-based layer a judge model sits on top of.