Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks A TACL 2026 paper by Shoumik Saha, Jifan Chen, Sam Mayers, Sanjay Krishna Gouda, Zijian Wang, and Varun Kumar introduces JAWS-BENCH, a benchmark that measures whether code-capable LLM agents actually compile and run malicious programs rather than merely refusing harmful text. Across seven LLM backends from five families, prompt-only attacks in the empty-workspace regime (JAWS-0) achieved 61% compliance, with 58% harmful, 52% parsing, and 27% running end-to-end; single-file (JAWS-1) compliance reached 100% for stronger models at a mean attack success rate of about 71%, and multi-file (JAWS-M) raised mean ASR to about 75% with 32% runnable attack code. Wrapping an LLM in an agent increased ASR by 1.6x by overturning initial refusals during planning and tool use, with similar trends in SWE-Agent and OpenAI Codex, motivating execution-aware defenses and refusal-preserving agent designs. Abstract Code-capable large language model LLM agents are embedded in software engineering workflows where they can read, write, and execute code, raising “jailbreak" stakes beyond text-only settings. Prior evaluations emphasize refusal or harmful-text detection, leaving open whether agents compile and run malicious programs. We present JAWS-BENCH Jailbreaks Across WorkSpaces , a benchmark spanning three escalating workspace regimes mirroring attacker capability: empty JAWS-0 , single-file JAWS-1 , and multi-file JAWS-M . We pair it with a hierarchical, executable-aware Judge Framework that tests i compliance, ii attack success, iii syntactic correctness, and iv runtime executability to measure de-ployable harm. Across seven LLM backends from five families, prompt-only attacks in JAWS-0 achieve 61% compliance; 58% are harmful, 52% parse, and 27% run end-to-end. In JAWS-1, compliance reaches 100% for stronger models with a mean ASR Attack Success Rate ≈ 71%; JAWS-M raises mean ASR to ≈ 75%, with 32% runnable attack code. Wrapping an LLM in an agent increases ASR by 1.6×, by overturning initial refusals during planning and tool use. Additional evaluations with SWE-Agent and OpenAI Codex exhibit similar trends, indicating that JAWS-BENCH can be reused across multiple agent frameworks. Category analyses identify which attack classes are most vulnerable and deployable, motivating execution-aware defenses and refusal-preserving agent designs. - Anthology ID: - 2026.tacl-1.94 - Volume: - Transactions of the Association for Computational Linguistics, Volume 14 https://aclanthology.org/volumes/2026.tacl-1/ - Month: - Year: - 2026 - Address: - Cambridge, MA - Venue: - TACL https://aclanthology.org/venues/tacl/ - SIG: - Publisher: - MIT Press - Note: - Pages: - 2081–2102 - Language: - URL: - https://aclanthology.org/2026.tacl-1.94/ https://aclanthology.org/2026.tacl-1.94/ - DOI: - 10.1162/tacl.a.792 https://doi.org/10.1162/tacl.a.792 - Cite ACL : - Shoumik Saha, Jifan Chen, Sam Mayers, Sanjay Krishna Gouda, Zijian Wang, and Varun Kumar. 2026. Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks https://aclanthology.org/2026.tacl-1.94/ . Transactions of the Association for Computational Linguistics , 14:2081–2102. - Cite Informal : - Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks https://aclanthology.org/2026.tacl-1.94/ Saha et al., TACL 2026 - PDF: - https://aclanthology.org/2026.tacl-1.94.pdf https://aclanthology.org/2026.tacl-1.94.pdf