arXiv:2609.19182v1 Announce Type: new Abstract: Benchmarks are central to how progress in large language models (LLMs) is assessed and communicated. Yet model rankings alone reveal little about how evaluation requirements themselves are changing. The expanding variety of benchmarks offers another perspective: what researchers expect LLMs to do, and what they count as successful performance. We systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026. Using staged screening and automated full-text coding, we examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms. The collection shows growing emphasis on action, interaction, and professional applications, while established and newer design elements frequently coexist. Model participation also develops unevenly: LLM-based scoring grows within both agent and non-agent groups, whereas model-generated materials show no comparable sustained increase in recent cohorts. These findings illuminate how public research translates capability expectations into concrete tests and criteria for success. As AI participates in constructing tests, performing tasks, and judging responses, they also raise a question: does expanding evaluation provide more independent evidence, or risk reproducing the preferences and blind spots of its participating models?
What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks
A systematic mapping of 14,767 arXiv papers introducing or updating LLM evaluation resources between January 2022 and August 2026 finds growing emphasis on action, interaction, and professional applications, with LLM-based scoring rising in both agent and non-agent groups while model-generated evaluation materials show no comparable sustained increase in recent cohorts. The study, arXiv:2609.19182v1, used staged screening and automated full-text coding to track changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms. The authors ask whether expanding evaluation provides more independent evidence or risks reproducing the preferences and blind spots of participating models.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.