The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance
A new study introducing BenchDrift shows that rephrasing benchmark problems while keeping meaning and answer fixed flips LLM correctness in both directions across eight models and three benchmarks (GS…