04:00
2026-08-03
arxiv.org
artificial-intelligence
Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation
A new arXiv preprint (2607.28801v1) introduces a dataset-centric meta-evaluation framework that audits LLM benchmark datasets at the sample level across five dimensions, revealing internal heterogenei…