When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation A systematic study of 60 language model benchmarks found that nearly half exhibit saturation, with rates increasing with age, and that resilience to saturation is impacted by expert-curation, not by public test data. The study, submitted by Mubashara Akhtar on arXiv on 18 Feb 2026 and revised through v3 on 29 Jun 2026, defines benchmark saturation and analyzes 14 properties to inform more durable evaluation approaches. Computer Science Artificial Intelligence Submitted on 18 Feb 2026 v1 https://arxiv.org/abs/2602.16763v1 , last revised 29 Jun 2026 this version, v3 Title:When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation View PDF /pdf/2602.16763 HTML experimental https://arxiv.org/html/2602.16763v3 Abstract:Artificial intelligence benchmarks are an important mechanism for measuring model progress and guiding deployment decisions. However, benchmarks quickly "saturate", making it difficult to differentiate models and diminishing their long-term value. In this study, we define benchmark saturation and analyze it across 60 language model benchmarks using 14 properties that relate to saturation. We find that nearly half of the our benchmarks exhibit saturation, with rates increasing with age. Further, we find that resilience to saturation is impacted by expert-curation, not by public test data. Our results suggest that design choices can extend benchmark longevity and inform more durable evaluation approaches. Submission history From: Mubashara Akhtar view email /show-email/a2e9c69b/2602.16763 Wed, 18 Feb 2026 16:51:37 UTC 222 KB v1 /abs/2602.16763v1 Sat, 30 May 2026 16:41:50 UTC 640 KB v2 /abs/2602.16763v2 v3 Mon, 29 Jun 2026 17:01:58 UTC 636 KB References & Citations Loading... Bibliographic and Citation Tools Bibliographic Explorer What is the Explorer? https://info.arxiv.org/labs/showcase.html arxiv-bibliographic-explorer Connected Papers What is Connected Papers? https://www.connectedpapers.com/about Litmaps What is Litmaps? https://www.litmaps.co/ scite Smart Citations What are Smart Citations? https://www.scite.ai/ Code, Data and Media Associated with this Article alphaXiv What is alphaXiv? https://alphaxiv.org/ CatalyzeX Code Finder for Papers What is CatalyzeX? https://www.catalyzex.com DagsHub What is DagsHub? https://dagshub.com/ Gotit.pub What is GotitPub? http://gotit.pub/faq Hugging Face What is Huggingface? https://huggingface.co/huggingface ScienceCast What is ScienceCast? https://sciencecast.org/welcome Demos Recommenders and Search Tools Influence Flower What are Influence Flowers? https://influencemap.cmlab.dev/ CORE Recommender What is CORE? https://core.ac.uk/services/recommender arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs https://info.arxiv.org/labs/index.html .