AI Evaluation Should Work With Humans
A position paper submitted to arXiv on 6 Jul 2026 argues that the dominant paradigm of AI evaluation, which focuses on superhuman autonomous performance, is guiding AI development in the wrong direction. The paper, title…