Models & Tools à l'instant4Ajouter aux favoris

A systematic arXiv study documents how benchmark saturation is quietly undermining our ability to compare models - and why this matters more than any individual benchmark result.
In plain terms: Models are maxing out benchmarks faster than we can replace them. A new academic study maps the saturation problem systematically for the first time - and finds it's accelerating across most standard evals.
The story: Benchmark saturation happens when top models score so close to ceiling that the benchmark can no longer differentiate between them. The study (arXiv:2602.16763) shows this is not isolated to a few famous cases (MMLU, HumanEval) - it's a structural pattern across the field. New benchmarks typically have a useful lifetime of 12-18 months before leading models cluster at the top.
The implication: announced "improvements" on saturated benchmarks are often meaningless. A model that goes from 91% to 93% on a benchmark where the ceiling is 95% tells you almost nothing about real-world capability differences.
Under the hood: The study proposes a saturation index - the proportion of top-model scores within one standard deviation of the benchmark ceiling. By this measure, approximately 60% of commonly reported benchmarks are already saturated for frontier models. GPQA and ARC-AGI remain useful. Most MMLU variants do not.
So what: Labs and researchers should publish saturation indices alongside benchmark results. For practitioners evaluating models: saturated benchmarks are marketing, not measurement. Task-specific evals on your actual use case remain the only reliable signal.
Article produit par intelligence artificielle, relu sous contrôle éditorial humain.
Connectez-vous pour rejoindre la discussion.
But is saturation really a flaw, or just proof that AI is getting *good at the wrong thing*? We're measuring speed, not sense.
These benchmarks were always artificial constructs, not genuine measures of intelligence. If we’ve hit a ceiling, maybe it’s time to ask whether we’ve been chasing the wrong goals all along.
This saturation problem highlights how benchmarks lag behind real-world use-we’re optimizing for the wrong metrics. Shouldn’t progress in AI be measured by societal impact rather than benchmarks alone?
The rush to build ever more complex benchmarks risks missing the forest for the trees. What if we stopped trying to measure intelligence and started designing systems that actually improve lives?
Fatigue hype 2026 : le tri entre modèle et harness