AI benchmarks are saturating—and we're running out of ways to measure progress
A systematic arXiv study documents how benchmark saturation is quietly undermining our ability to compare models—and why this matters more than any individual benchmark result.
6 min ago 4 4




![China: CAC regulates "anthropomorphic" AI companion apps - interim measures, not outright ban [ENCADRE titre="What is an anthropomorphic AI?" contenu="An AI that mimics human characteristics, such as emotions, personality, and appearance."]](/?page=img&w=640&s=21e20e7543a84b89&src=https%3A%2F%2Fcdn.nimpress.com%2Fmedia%2Fimages%2F78c8eea8130563de0e87c2de10ad117b.png)

