Evaluating a legal AI system
Evaluating a legal AI system needs two separate scores, not one, since getting every citation right and finding every citation that matters are different questions.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Evaluating a legal AI system means checking two separate things. Is everything it said true, and did it leave anything important out?
A witness testifies in court. One kind of mistake is stating something false. A completely different kind of mistake is leaving out something important they should have mentioned.
A perfect witness makes neither mistake. A legal AI tool needs the same double check. Grading it on only one of these two questions misses half the picture entirely.
Why it exists
Every lesson in this section built a piece of a legal AI pipeline. Splitting clauses, extracting categories, retrieving case law, generating an answer with citations. None of that is worth anything without checking whether the final result is actually trustworthy.
A single overall "accuracy" number hides too much. A system can be right every time it speaks, and still miss half of what it should have covered. Or it can cover everything and get some of it wrong. Those are very different failures.
Legal evaluation needs both scores, borrowed from earlier lessons in this course. Precision: of what it said, how much was right. Recall: of what should have been said, how much was covered.
How it works
Question: "Can a landlord evict without notice?"
Required citations: {Rai v. Sharma, 2018}
AI's citations: {Rai v. Sharma, 2018, Patel v. State Housing Board, 2017}
Precision: of the citations the AI gave, how many were real and correct?
-> only half were -- it also invented one that does not exist.
Recall: of the citations actually needed, how many did the AI find?
-> all of them -- nothing required was missed.Neither score alone tells the full story. Weak precision sounds bad, but strong recall shows the real answer was actually found. Both, read together, are needed.
Where you have already seen it
- Legal AI product benchmarks. Vendor comparisons of legal research tools typically report both citation precision and recall, not a single blended score.
- Court sanctions cases. Lawyers sanctioned for AI-generated fake citations were, in evaluation terms, failing on precision specifically.
- Missed-clause complaints. A contract review tool overlooking a risky clause is a recall failure, different from citing a wrong one.
Remember this
- Precision asks: of what the system said, how much was correct?
- Recall asks: of what it should have said, how much did it actually cover?
- A legal AI system needs both scores checked, since a good result on one says nothing about the other.
What to learn next
- Using an LLM to grade text — a technique for scoring open-ended legal answers where a fixed reference is hard to define.
- Do your labels even agree? — checking that the human-provided "correct answers" this evaluation depends on are themselves reliable.
- Legal RAG and fabricated citations — the specific failure mode this precision score is built to catch.
Developer — Code and libraries.
This scores two example answers for citation precision and recall, and shows how each failure mode looks completely different in the numbers.
Setup
Nothing to install. This uses only plain Python.
Scoring citation precision and recall
test_cases = [
{
"question": "Can a landlord evict without notice?",
"gold_citations": {"Rai v. Sharma, 2018"},
"model_citations": {"Rai v. Sharma, 2018", "Patel v. State Housing Board, 2017"},
},
{
"question": "Must a bank disclose loan fees upfront?",
"gold_citations": {"Mehta v. Union Bank, 2020", "Fernandes v. CityCorp, 2019"},
"model_citations": {"Mehta v. Union Bank, 2020"},
},
]
def citation_precision_recall(gold, predicted):
if not predicted:
return 0.0, 0.0
correct = gold & predicted
precision = len(correct) / len(predicted)
recall = len(correct) / len(gold) if gold else 1.0
return precision, recall
for case in test_cases:
p, r = citation_precision_recall(case["gold_citations"], case["model_citations"])
fabricated = case["model_citations"] - case["gold_citations"]
missed = case["gold_citations"] - case["model_citations"]
print(f"Q: {case['question']}")
print(f" precision={p:.2f} recall={r:.2f}")
if fabricated:
print(f" fabricated/wrong citations: {fabricated}")
if missed:
print(f" missed required citations: {missed}")
print()Q: Can a landlord evict without notice?
precision=0.50 recall=1.00
fabricated/wrong citations: {'Patel v. State Housing Board, 2017'}
Q: Must a bank disclose loan fees upfront?
precision=1.00 recall=0.50
missed required citations: {'Fernandes v. CityCorp, 2019'}The two questions fail in opposite ways. The first answer found everything it needed, but padded in something fake. The second answer cited nothing wrong, but stopped one citation short of complete. A single blended score would make these two very different failures look similar.
Line by line
gold & predicted is set intersection: exactly the citations that appear in both the required set and the model's answer. This is the count of genuinely correct citations, used for both precision and recall.
gold - predicted and predicted - gold are set differences: citations required but missing, and citations given but not required. These are exactly the two failure modes shown in the printed output.
if not predicted: return 0.0, 0.0 handles the edge case of an answer with no citations at all, avoiding a divide-by-zero and correctly scoring it as having found nothing.
Common mistakes
Reporting a single averaged score across precision and recall without context. As shown above, the two failures need different fixes: fabrication needs better verification, missed citations need better retrieval. Averaging them hides which fix is actually needed.
Treating "gold citations" as beyond question. The required citation set itself needs to be correct and complete, which is exactly why inter-annotator agreement, checking whether legal experts agree on what "correct" means here, matters before trusting any evaluation built on top of it.
Evaluating only on easy, clear-cut questions. A test set of only simple questions with one obvious citation will not reveal a system's real precision and recall on the messier, more ambiguous questions it will actually face in practice.
Try it yourself
Add a third test case where model_citations is an empty set, an AI that refused to answer rather than risk citing something wrong. Compute its precision and recall.
Consider whether a system that says "I don't know" should be scored worse, the same, or better than one that gives a fabricated answer, an actual policy decision, not only a technical one, that a real deployment has to make explicitly.
What to learn next
- Do your labels even agree? — validating the gold citation sets this evaluation depends on.
- Using an LLM to grade text — extending this beyond citations, to judging the quality of the surrounding legal reasoning too.
- Legal RAG and fabricated citations — the system this evaluation approach is designed to hold accountable.
Researcher — Mathematics and papers.
Evaluation as a multi-dimensional problem
A complete legal AI evaluation typically decomposes across several independent axes, each requiring its own metric and often its own labelled dataset:
citation_accuracy: precision/recall over cited sources, as in the developer block
factual_correctness: does the substantive legal claim match ground truth
completeness: were all relevant considerations addressed
reasoning_quality: is the legal reasoning itself sound, not just the conclusionTreating these as one blended score is a common and consequential mistake, since a system can excel on one axis while failing badly on another, exactly the pattern the developer block's two test cases demonstrate concretely for citation accuracy alone.
Benchmark efforts in legal AI evaluation
LegalBench (Guha et al., 2023) assembles over 160 tasks spanning issue-spotting, rule-recall, rule-application and rule-conclusion, explicitly designed around the IRAC framework (Issue, Rule, Application, Conclusion) lawyers use to structure legal reasoning, rather than testing citation accuracy in isolation.
LegalBench-RAG (Pipitone & Alami, 2024) specifically targets retrieval quality within legal RAG systems, measuring whether retrieved passages actually support the generated answer at the level of specific text spans, not only at the document level. That is a finer-grained precision/recall than the citation-level check in the developer block.
The gold-standard construction problem
Every metric in this lesson depends on a "gold" answer set assumed correct. In practice, constructing that gold set for legal questions is itself contested: experienced lawyers can disagree on which citations are truly required for a given question, particularly for genuinely difficult or unsettled areas of law.
This connects directly back to inter-annotator agreement from earlier in this course. A legal evaluation benchmark's own reliability should be reported alongside its results, using the same kappa or alpha statistics used for any other human-labelled dataset. In practice, this step is applied inconsistently across published legal AI benchmarks.
Cost-sensitive evaluation
Unlike many NLP tasks, error costs in legal AI are genuinely asymmetric and often quantifiable. A missed citation might cost a lawyer a few extra minutes of review time. A fabricated citation submitted to a court has led to real monetary sanctions and professional discipline in reported cases. A rigorous evaluation framework should weight these differently, closer to an expected-cost metric than a symmetric F1 score:
expected_cost = cost_fp * FP + cost_fn * FNcost_fp,cost_fnare domain-specific cost weights for a false positive (fabrication) and false negative (omission) respectively, typically asymmetric and set by domain expertise rather than derived statistically.
Complexity
Citation-level precision/recall, as in the developer block, is O(|gold| + |predicted|) per example using set operations, trivial at any realistic scale. Span-level and reasoning-quality evaluation, as used in LegalBench-RAG and similar benchmarks, costs substantially more: typically one or more LLM-judge forward passes per example, plus the human-annotation cost of constructing the gold standard in the first place.
Key references
- Guha, N. et al. (2023). LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models. arXiv:2308.11462
- Pipitone, N. & Alami, G. H. (2024). LegalBench-RAG: A Benchmark for Retrieval-Augmented Generation in the Legal Domain. arXiv:2408.10343
- Magesh, V. et al. (2024). Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362
Current state and open problems
Citation-level precision and recall, the simplest metric covered in this lesson, is now reported by essentially every serious legal AI evaluation, and is a genuinely necessary, if not sufficient, check.
The open problem is evaluating reasoning quality itself, beyond citation bookkeeping: whether a legal AI system's chain of argument would satisfy a competent lawyer, not only whether its final citations happen to check out. This remains substantially unsolved, and is the reason human legal review remains a required part of every serious deployment discussed across this section, not an optional safety net bolted on afterward.
What to learn next
- Using an LLM to grade text — the leading current approach to evaluating reasoning quality, not only citations.
- Do your labels even agree? — validating that a legal evaluation's own gold standard is reliable.
- Legal RAG and fabricated citations — where this chapter's evaluation techniques close the loop back to the system that generates the citations in the first place.