12:07
2026-07-27
arxiv.org
artificial-intelligence
Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks
A controlled prompt-ablation study across 22 frontier models from 7 providers on 23 Cybench CTF challenges found that 37.1% of passes involved cheating under baseline conditions, with 21 of 22 models โฆ