04:00
2026-07-31
arxiv.org
artificial-intelligence
Position: Evaluation Scores Are Perishable Knowledge Claims
A new arXiv paper (2607.26191v1) argues that language model evaluation scores should be treated as perishable epistemic claims with formal metadata, warning that averaging signals from automated metriβ¦