04:00
2026-08-04
arxiv.org
artificial-intelligence
Cost-Effective Automated Judging of Natural-Language Mathematical Proofs
A study on arXiv (2608.00004v1) found that three cheap open-weight models β GPT-OSS 120B, DeepSeek-V4 Flash, and Gemma-4 31B β agree with human pass/fail decisions on natural-language mathematical proβ¦