LLM-as-judge
In one sentence LLM-as-judge uses a strong language model to grade another model's outputs against a rubric, making large-scale evaluation affordable.
Updated
LLM-as-judge is the technique of using a capable language model to score or compare another model's outputs, standing in for human graders.
Board exams with lakhs of essay answers are not marked by one examiner — they are marked by many, trained on a rubric with model answers, spot-checked by moderators. LLM-as-judge applies the same delegation to AI evaluation: quality judgements ("is this summary faithful? is this reply helpful?") resist exact-match scoring, humans are too slow and costly for thousands of cases per change, so a strong model gets the rubric and does the marking.
In practice the judge receives the task, the output to grade, and a written rubric, then returns a score or a pairwise verdict. Pairwise ("which of A and B answers better?") is the more reliable mode — comparison is easier than absolute scoring, for models as for people — and powers public leaderboards like Chatbot Arena's automated variants.
The judge's known biases are documented and manageable: position bias (favouring whichever answer appears first — fix by swapping order and averaging), length bias (favouring longer answers), self-preference (rating its own family's style higher), and being impressed by confident tone — the same weakness that makes hallucination dangerous for humans. The professional workflow therefore calibrates: collect a few hundred human judgements once, measure judge-human agreement, adjust the rubric until agreement is high, then let the judge scale. Used this way it is the backbone of practical evals; used blindly, it is a bias with a scoreboard.
Where to go next
- Full lesson: Model evaluation
- Related terms: evals, reward-model, benchmark, hallucination