Member-only story
plain-English guide to what AI benchmark scores like MMLU, GPQA, SWE-bench, ARC-AGI, and Arena Elo actually measure, why the differences between top models are often noise, and how to read a benchmark claim without getting fooled.
A friend sent me a screenshot last week: a new model launch, a bar chart, five benchmark names she’d never heard of, and a caption claiming “state of the art.” Her question was simple and completely fair: does this number mean the model is smart, or does it mean the company that made it is good at picking favorable comparisons?
Both, usually, and telling them apart is the actual skill here. Benchmark scores aren’t lies, but they’re also not the clean report card the marketing implies. A model can genuinely lead one benchmark and lose badly on the exact skill you care about, and both facts can be true on the same launch day. This guide walks through what these numbers actually measure, why the gap between two impressive-sounding scores is often statistical noise, and how to read a claim like an evaluator instead of a fan.
The short version #
A benchmark score tells you how a model performed on one specific, fixed set of tasks, scored one specific way, often using an evaluation setup the model’s own creator chose. That’s it…