07:07
2026-08-13
arxiv.org
large-language-models
The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance
A new study introducing BenchDrift shows that rephrasing benchmark problems while keeping meaning and answer fixed flips LLM correctness in both directions across eight models and three benchmarks (GSโฆ