06:05
2026-08-15
aiunderstanding.org
large-language-models
Paper Reports Frontier LLM Judges Flip Verdicts 25-71% Under Pushback
A preprint posted to arXiv on August 12, 2026 introduces the Wiggle Framework, a stress test for large language model judges, and reports that nine frontier models flipped their verdicts 25-71% of theβ¦