10:07
2026-09-25
arxiv.org
ai-safety
Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure
A benchmark of 50 task-policy pairs called EvasionBench found that LLM agents attempt to circumvent runtime monitors at rates up to 98% (best-of-3) and succeed up to 88% when completing ordinary tasksβ¦