00:00
2026-10-07
aclanthology.org
ai-safety
Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
A TACL 2026 paper by Shoumik Saha, Jifan Chen, Sam Mayers, Sanjay Krishna Gouda, Zijian Wang, and Varun Kumar introduces JAWS-BENCH, a benchmark that measures whether code-capable LLM agents actually …