AI glossary

Evals

In one sentence Evals are systematic tests for LLM behaviour — your own graded question set, run on every change, because "it seems better" is not measurement.

By Updated

Evals are structured, repeatable tests that measure how well an LLM or LLM product performs on the tasks you actually care about.

They are the tasting menu a restaurant runs before changing suppliers. The head chef does not switch rice brands because the new one "seems fine" — she cooks the same ten dishes with both and compares. LLM development needs the same discipline, because every change — a new model version, a reworded prompt, different RAG chunking — shifts behaviour in ways eyeballing three examples will not reveal. The standing failure mode of LLM teams is vibes-driven iteration: change something, try two prompts, ship, regress silently.

An eval is concretely three things: a dataset of test cases (real user queries, edge cases, past failures), a grading method, and a harness that runs the set and reports scores per change. Grading is the craft:

exact / programmatic  answer matches, JSON parses, code passes tests   cheap, objective
LLM-as-judge          a strong model grades against a rubric           scales, needs calibration
human review          the gold standard                                slow, used to spot-check

Public benchmarks rank models in general; evals rank your system on your task — the ranking that pays your bills, and the two disagree often. Practical habits that compound: start with twenty cases rather than waiting for five hundred; grow the set from production failures (ground-truth you paid for); run evals in CI like unit tests; and keep a slice for safety regressions from red-teaming.

Where to go next