Evaluating agents
How to measure an agent honestly — success rates over many runs, trajectory review, the caveats of AI judges, and regression suites that catch backsliding.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Evaluating an agent means running it many times on real tasks and counting, honestly, what happened.
Think of a driving test. The examiner does not check only whether the car reached the destination. They watch the whole drive — mirrors, signals, speed, that one moment at the roundabout.
And nobody licenses a driver off one lucky Sunday trip. You would want several drives, on different roads, on a bad-traffic day too.
Why it exists
An agent — a model that takes many steps and actions to finish a job, see AI agents — fails in ways a single-answer chatbot does not.
It is different every run. Models have randomness built in — temperature and sampling explains why. The demo that worked in the meeting can fail an hour later, unchanged.
Small slips stack up. A worker who gets each task right nine times out of ten will still spoil most long jobs, because a long job needs every step right in a row.
The destination hides the journey. An agent can produce the right answer after a wasteful, alarming detour — or fail politely one step from success. The final answer alone tells you almost nothing.
How it works
same task, twenty runs:
✓ ✓ ✓ ✗ ✓ ✓ ✓ ✓ ✗ ✓ ✓ ✓ ✓ ✓ ✗ ✓ ✓ ✓ ✓ ✓
| |
v v
count them open the failures
→ success rate → read every step the agent took
(its "trajectory") and find where
it went wrongFour habits make up honest agent evaluation:
- Success rate, not success story. Run each task many times; report the fraction that passed.
- Trajectory review. The trajectory is the full list of steps the agent took. Read the failures like an examiner reads the drive.
- A careful eye on AI examiners. Teams often ask a second model to grade the first — an LLM judge. Useful, but it is like asking one student to grade another: fine for neat handwriting, risky for whether the answer is right.
- A regression suite. A fixed set of tasks re-run after every change — the same driving circuit, driven again after every repair — so you notice when a "fix" breaks something older. A regression is exactly that: something that used to work and quietly stopped.
Where you have already seen it
Board exams moderate marks across examiners. Cricket selectors watch months of matches, not one good innings. And the team behind your bank's chatbot replays hundreds of old customer conversations against every new version before releasing it.
The honest part
This is the least glamorous work in the entire agents field, and the most valuable. Demos are cheap; run-it-two-hundred-times tables are rare. Teams that skip this ship agents that work in the meeting and fail with customers. There is no clever trick coming to replace the counting.
Remember this
- One successful demo is not evidence. Twenty runs are the beginning of evidence.
- Read trajectories, not only final answers — the failure lives in the steps.
- Re-run a fixed task suite after every change, or improvements will quietly break old wins.
What to learn next
- Orchestrating multiple agents — where evaluation gets harder still.
- Model evaluation — the classical foundations this builds on.
- Temperature and sampling — why identical runs differ at all.
Developer — Code and libraries.
Setup
python3 --version # standard library only — nothing to installA minimal harness you can keep
The stub below stands in for your agent so the page prints identically everywhere. The harness around it — trials, rates, the reliability column — is the part you keep.
import random
random.seed(7) # fixed seed, so this page and your machine print the same table
# Stand-in agent: succeeds with a hidden probability per task.
# Replace this one function with a call to your real agent plus its checker.
TRUE_RATES = {"refund flow": 0.9, "address change": 0.6, "cancel order": 0.3}
def agent_passes(task):
return random.random() < TRUE_RATES[task]
TRIALS = 20
print(f"{'task':<16}{'pass rate':>10}{'est. 8-in-a-row':>18}")
for task in TRUE_RATES:
results = [agent_passes(task) for _ in range(TRIALS)]
rate = sum(results) / TRIALS
print(f"{task:<16}{rate:>9.0%}{rate ** 8:>17.1%}")task pass rate est. 8-in-a-row refund flow 95% 66.3% address change 75% 10.0% cancel order 30% 0.0%
The third column is the lesson
A 95% agent sounds finished. But a customer who needs it to work eight times in a row gets a clean streak only 66% of the time. The 75% agent almost never manages a week without a failure.
The two framings have names. pass@k asks: did at least one of k attempts succeed? It flatters, and suits situations where retries are cheap. pass^k asks: did all k attempts succeed? It describes what a user relying on the agent actually experiences. Report pass^k for anything customer-facing.
Note also what the pass rates did: the measured 95% came from a true 90% coin, and 75% from a true 60%. Twenty trials carry real statistical noise — enough to misrank two similar agents. More trials, and error bars, are not optional extras.
What "passes" must mean
The checker deserves more thought than the agent:
- Check the world, not the words. After "cancel order", assert the database row says
cancelled. An agent that says "Done!" without doing anything must fail the check. Wording checks pass liars. - Check the trajectory too. Assert the forbidden tool was never called, the step count stayed under a cap, no action ran after the task was done. These catch the agent that succeeds alarmingly.
- Reset the environment between trials. Trial two must not inherit trial one's cancelled order, or your rates measure contamination, not competence.
LLM-as-judge, with its caveats attached
For outputs with no programmatic check — tone, helpfulness, summary quality — a judge model scoring against a written rubric is the standard tool. Use it with its known biases in view:
- Position bias: in A-vs-B comparisons, judges favour whichever answer is presented first. Score both orders and average.
- Verbosity bias: longer answers score higher at equal quality.
- Self-preference: models rate their own family's style generously. Judge with a different model than the one that acted.
- Weak at arithmetic and facts. Anything checkable by code should be checked by code, never by judge.
Calibrate before trusting: have a human label fifty transcripts, measure agreement with the judge, and re-measure whenever the judge model or rubric changes.
The regression suite
Collect every real failure you fix into a task list with checkers. Re-run the whole list on every prompt edit, tool change, or model upgrade, and compare rates against the last release. This is the agent version of the discipline in model evaluation — and it is how you stop the Tuesday fix from breaking the Monday feature. Grow it from production failures; they are the tasks reality has voted for.
Common mistakes
Evaluating on one run per task. The single biggest one. Randomness makes one run a coin flip, not a measurement.
Judging with the actor. The same model (or family) both doing and grading inflates scores through self-preference.
Changing the task set every week. Rates stop being comparable, and trends — the entire point — vanish.
Ignoring cost and steps. An agent that succeeds in 40 steps and 200 rupees of tokens per task can be worse than one that fails fast. Record steps, tokens, and latency alongside pass rates.
Try it yourself
Change the seed from 7 to 8 and rerun. The rates move — that movement is your uncertainty, made visible. Then add a steps count to the stub (say, a random 3 to 12) and report mean steps per task next to the pass rate. You now have a two-metric harness, which is one metric more than most teams run.
What to learn next
- Orchestrating multiple agents — where evaluation gets harder still.
- Model evaluation — the classical foundations this builds on.
- Temperature and sampling — why identical runs differ at all.
Researcher — Mathematics and papers.
Reliability formalised
For per-trial success probability $r$ and $k$ independent trials: $\text{pass@}k = 1 - (1-r)^k$ and $\text{pass}^k = r^k$. Symbols: $r$ the true task-level success rate, $k$ the trial count. The independence assumption is generous — trials share the model, prompts, and often environment state, and correlated failures make realised $\text{pass}^k$ worse than $r^k$. Yao et al. (2024), τ-bench (arxiv.org/abs/2406.12045), introduced pass^k for tool-agent-user settings and showed it falling steeply in $k$ even for frontier models — the headline argument that agents were far less reliable than single-run leaderboards implied.
Estimation error: a measured rate $\hat{r}$ over $n$ trials has standard error $\sqrt{\hat{r}(1-\hat{r})/n}$; at $\hat{r} = 0.5, n = 20$ the 95% interval spans roughly $\pm 22$ points. Use the Wilson interval near 0 or 1 rather than the normal approximation, and treat two agents whose intervals overlap as unranked. Distinguish seed variance from task variance: report per-task rates, not only the aggregate, or a regression on one task hides inside an average.
Benchmarks
- Mialon et al. (2023), GAIA (arxiv.org/abs/2311.12983) — tool-use questions; humans 92%, GPT-4 with plugins 15% at publication. Scaffolding, not model size, drives most of the spread.
- Liu et al. (2023), AgentBench (arxiv.org/abs/2308.03688) — eight interactive environments, from OS shells to databases.
- Zhou et al. (2023), WebArena (arxiv.org/abs/2307.13854) — realistic self-hosted web tasks with programmatic checkers.
- Jimenez et al. (2023), SWE-bench (arxiv.org/abs/2310.06770) — real GitHub issues; resolution verified by held-out tests. The de facto coding-agent benchmark, and a cautionary tale in contamination management.
- The Berkeley Function Calling Leaderboard scores single tool calls by AST match — a unit test to the benchmarks' integration tests.
Public benchmarks are upper-bound advertising for your use case: contamination (tasks leaking into training data), scaffold overfitting, and domain mismatch all bias upward. A private, versioned suite drawn from your own traffic is the measurement that pays rent.
LLM-as-judge, quantified
Zheng et al. (2023), Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (arxiv.org/abs/2306.05685), document position bias, verbosity bias, self-enhancement bias, and weak math grading, while reporting that a strong judge agrees with human raters at above 80% on their data — comparable to human-human agreement there. The transferable protocol, not the number: measure judge-human agreement on your distribution before the judge's scores mean anything. For agent trajectories specifically, judges inherit an extra failure: they grade the transcript's storytelling, and a confident narration of a wrong trajectory reads well. Anchor the judge with programmatic trajectory facts (tools called, states reached) surfaced into its rubric.
Failure taxonomy
Cemri et al. (2025), Why Do Multi-Agent LLM Systems Fail? (arxiv.org/abs/2503.13657), annotate 1,600+ traces into 14 failure modes across three classes — system design, inter-agent misalignment, and task verification. Verification failures (weak or missing checkers) are a class of their own, which is the research-grade restatement of this lesson: the checker is a first-class component, and most measured "agent" failures include an evaluation failure.
What to learn next
- Orchestrating multiple agents — where evaluation gets harder still.
- Model evaluation — the classical foundations this builds on.
- Temperature and sampling — why identical runs differ at all.