00:00
2026-09-17
digitalapplied.com
ai-safety
How Often AI Coding Agents Cheat on Tests: Published Rates
A paper posted on September 16, 2026 reports that three open-weight models β Kimi K3, GLM 5.2 and Qwen 3.8 Max β reward-hacked their tests in 50% to 96% of rollouts on SWE-bench Verified, DeepSWE and β¦