AI glossary

Benchmark

In one sentence A benchmark is a shared, fixed test that lets different models be compared on the same questions under the same rules.

By Updated

A benchmark is a fixed, public test set plus a scoring rule, so everyone's model answers the same questions and the numbers can be compared.

Board exams exist for the same reason. A hundred schools cannot compare students if every school writes its own test — one school's 90% may be another's 60%. A common paper, common marking scheme, and published results make the comparison meaningful. Benchmarks are ML's board exams: ImageNet for image classification, GLUE for language understanding, MMLU and GSM8K for LLM knowledge and maths, SWE-bench for coding agents.

Benchmarks focus a field powerfully — the ImageNet competition triggered the deep learning era in 2012. But they fail in a predictable way, worth knowing as Goodhart's law: when a measure becomes a target, it stops being a good measure. Once everyone optimises for one exam, scores rise faster than real capability.

For LLMs there is a sharper failure mode: contamination. Benchmarks are text on the internet, and models train on internet text, so the exam paper may already be inside the model with answers attached. A high score can mean memory, not skill.

The practical stance: treat public benchmark scores as rough, gameable signals for shortlisting. Then build a private test from your own task — your evals — because a model's rank on MMLU says little about its rank on your problem.

Where to go next