A Study of 60 AI Benchmarks Found Nearly Half Have Saturated — Which Is Why 'State of the Art' Keeps Meaning Less
When the top models bunch near a benchmark's ceiling, the score stops telling you which one is better. A 37-author study finds 29 of 60 benchmarks have saturated — and that how hard the field optimises against a test matters more than its age.

Every few weeks a lab announces a new model that "sets a new state of the art," and the number it points at is a benchmark score. A new paper looking at 60 of those benchmarks explains why that phrase is worth less than it used to be: nearly half of the tests have worn out.
What "saturation" means
A benchmark saturates when the best models bunch up near the top of it. Once the leaders are all scoring 92, 93, 94 per cent, the test can no longer tell you which model is actually better — the differences fall inside the noise. The benchmark still produces a number, but the number has stopped carrying information. It has become a ceiling rather than a ruler.
The study — When AI Benchmarks Plateau, a 37-author systematic review first posted in February and revised through June — measured this directly across 60 language-model benchmarks. Its headline finding: 29 of the 60 (48 per cent) show high or very high saturation, and 14 of those are in the "very high" band, where the top models are essentially tied. Named examples give the shape of it — an older coding test like HumanEval sits among the saturated (the paper points to ImageNet in vision as the same story playing out beyond language), while harder, more recent designs like ARC-AGI and BIG-Bench Hard remain resilient.
The part that resists the easy story
The tidy explanation would be "benchmarks just get old and everyone catches up." The paper does find that older benchmarks are more saturated on average — mean saturation rises from 0.51 for tests under two years old to 0.60 for those over five — but it is careful to say this trend is modest and not statistically significant at conventional thresholds. Age alone is not the driver.
What drives it, in the authors' words, is "structural exposure dynamics and measurement resolution limits, rather than by isolated design choices." In plainer terms: a benchmark wears out faster the more the field optimises against it, and it stops discriminating once its test set is too small to separate models that are closely matched. A benchmark can be young and already saturated if everyone is training toward it; it can be old and still sharp if it was built well.
That "built well" turns out to matter. Expert-curated benchmarks showed lower saturation at comparable ages, and several stayed unsaturated despite years of exposure, while crowdsourced sets wore out faster. Careful construction buys a test a longer working life.
Why it matters
This is the measured version of a suspicion this masthead keeps returning to: a leaderboard number is a claim, not a verdict. When a model tops a saturated benchmark, it has cleared a bar that most of its rivals also clear — which tells you it is competent, not that it is best. It is the same caution that applies to a lab reporting its own scores on public datasets, as Mistral did this week with its Shieldstral safety model: the dataset can be standard and the result still uninformative if everyone is already near the top of it.
The authors' recommendations are the practical upshot, and they read as a maintenance schedule rather than a redesign: use bigger or stratified test sets so there is room to separate leaders; refresh benchmarks with new or adversarial items so optimisation cannot fully catch up; report confidence intervals so a one-point lead is not mistaken for a real one; and set explicit criteria for when a benchmark should be retired.
The takeaway
Benchmarks are consumables with a shelf life, not fixed yardsticks — and the study puts a number on how many are past their best. It does not mean the field cannot measure progress; it means the honest measure is moving, which is why the hardest current tests are new ones built specifically to resist saturation. The next time a model "tops the charts," the useful question is not how high it scored, but whether the chart it topped can still tell one strong model from another.
Ask Relay — he reads every question himself and replies personally by email.
