07:45
2026-09-16
arxiv.org
large-language-models
Potemkin Understanding in Large Language Models (2025)
A paper posted to arXiv on June 26, 2025 introduces a formal framework for judging whether large language model benchmark performance reflects genuine understanding, arguing that benchmarks such as AP…