04:00
2026-07-20
arxiv.org
artificial-intelligence
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning
A new study on arXiv (2607.15388v1) testing 4,181 Omni-MATH problems with gpt-oss-120b actors finds that broadcast-style peer discussion achieves higher final accuracy than a planner-executor-reviewer…