Benchmarks Are Not Progress

A number that improves every quarter is reassuring and largely decorative. Measuring what a system is good at is much harder than ranking it.

Benchmark scores are the currency of AI progress reporting, for the understandable reason that they are the only comparable numbers available. They also mislead in ways that are well known to the people producing them and rarely mentioned in the coverage.

Contamination

Benchmarks are published on the internet. Models are trained on the internet. A test set that has leaked into training data measures memorisation, and nobody outside the training organisation can verify whether it has. This is not a hypothetical concern; it is a structural feature of how these systems are built.

Optimising for the measure

When a benchmark becomes the way capability is reported, effort flows towards it. That is not cheating, it is what measurement does to any field. It does mean that a rising score can reflect increasing attention to the test rather than increasing general capability — and the two are indistinguishable from the outside.

The tasks are not the work

Multiple-choice questions and self-contained coding puzzles are chosen because they can be graded automatically, not because they resemble what anyone actually does. Real tasks are ambiguous, involve incomplete information, span long contexts, and are judged by whether a person was helped. Almost none of that is captured by anything on a leaderboard.

What would be better

Held-out evaluations nobody can train on. Measurement on tasks that people genuinely perform, graded by the people who perform them. Reporting the distribution of outcomes rather than an average, since the tail is where harm lives. And, for anyone deploying, an evaluation set built from their own data — which will tell them more than every public benchmark combined.

Get the next one by email