{"slug": "breaking-the-code-security-assessment-of-ai-code-agents-through-systematic", "title": "Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks", "summary": "A TACL 2026 paper by Shoumik Saha, Jifan Chen, Sam Mayers, Sanjay Krishna Gouda, Zijian Wang, and Varun Kumar introduces JAWS-BENCH, a benchmark that measures whether code-capable LLM agents actually compile and run malicious programs rather than merely refusing harmful text. Across seven LLM backends from five families, prompt-only attacks in the empty-workspace regime (JAWS-0) achieved 61% compliance, with 58% harmful, 52% parsing, and 27% running end-to-end; single-file (JAWS-1) compliance reached 100% for stronger models at a mean attack success rate of about 71%, and multi-file (JAWS-M) raised mean ASR to about 75% with 32% runnable attack code. Wrapping an LLM in an agent increased ASR by 1.6x by overturning initial refusals during planning and tool use, with similar trends in SWE-Agent and OpenAI Codex, motivating execution-aware defenses and refusal-preserving agent designs.", "body_md": "##### Abstract\n\nCode-capable large language model (LLM) agents are embedded in software engineering workflows where they can read, write, and execute code, raising “jailbreak\" stakes beyond text-only settings. Prior evaluations emphasize refusal or harmful-text detection, leaving open whether agents compile and run malicious programs. We present JAWS-BENCH(Jailbreaks Across WorkSpaces), a benchmark spanning three escalating workspace regimes mirroring attacker capability: empty (JAWS-0), single-file (JAWS-1), and multi-file (JAWS-M). We pair it with a hierarchical, executable-aware Judge Framework that tests (i) compliance, (ii) attack success, (iii) syntactic correctness, and (iv) runtime executability to measure de-ployable harm. Across seven LLM backends from five families, prompt-only attacks in JAWS-0 achieve 61% compliance; 58% are harmful, 52% parse, and 27% run end-to-end. In JAWS-1, compliance reaches 100% for stronger models with a mean ASR (Attack Success Rate) ≈ 71%; JAWS-M raises mean ASR to ≈ 75%, with 32% runnable attack code. Wrapping an LLM in an agent increases ASR by 1.6×, by overturning initial refusals during planning and tool use. Additional evaluations with SWE-Agent and OpenAI Codex exhibit similar trends, indicating that JAWS-BENCH can be reused across multiple agent frameworks. Category analyses identify which attack classes are most vulnerable and deployable, motivating execution-aware defenses and refusal-preserving agent designs.\n- Anthology ID:\n- 2026.tacl-1.94\n- Volume:\n- [Transactions of the Association for Computational Linguistics, Volume 14](https://aclanthology.org/volumes/2026.tacl-1/)\n- Month:\n- Year:\n- 2026\n- Address:\n- Cambridge, MA\n- Venue:\n- [TACL](https://aclanthology.org/venues/tacl/)\n- SIG:\n- Publisher:\n- MIT Press\n- Note:\n- Pages:\n- 2081–2102\n- Language:\n- URL:\n- [https://aclanthology.org/2026.tacl-1.94/](https://aclanthology.org/2026.tacl-1.94/)\n- DOI:\n- [10.1162/tacl.a.792](https://doi.org/10.1162/tacl.a.792)\n- Cite (ACL):\n- Shoumik Saha, Jifan Chen, Sam Mayers, Sanjay Krishna Gouda, Zijian Wang, and Varun Kumar. 2026. [Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks](https://aclanthology.org/2026.tacl-1.94/) .*Transactions of the Association for Computational Linguistics* , 14:2081–2102.\n- Cite (Informal):\n- [Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks](https://aclanthology.org/2026.tacl-1.94/) (Saha et al., TACL 2026)\n- PDF:\n- [https://aclanthology.org/2026.tacl-1.94.pdf](https://aclanthology.org/2026.tacl-1.94.pdf)", "url": "https://wpnews.pro/news/breaking-the-code-security-assessment-of-ai-code-agents-through-systematic", "canonical_source": "https://aclanthology.org/2026.tacl-1.94/", "published_at": "2026-10-07 00:00:00+00:00", "updated_at": "2026-10-08 00:18:07.865387+00:00", "lang": "en", "topics": ["ai-safety", "ai-agents", "large-language-models", "ai-research", "artificial-intelligence"], "entities": ["JAWS-BENCH", "Shoumik Saha", "Jifan Chen", "Sam Mayers", "Sanjay Krishna Gouda", "Zijian Wang", "Varun Kumar", "Transactions of the Association for Computational Linguistics"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/breaking-the-code-security-assessment-of-ai-code-agents-through-systematic", "markdown": "https://wpnews.pro/news/breaking-the-code-security-assessment-of-ai-code-agents-through-systematic.md", "text": "https://wpnews.pro/news/breaking-the-code-security-assessment-of-ai-code-agents-through-systematic.txt", "jsonld": "https://wpnews.pro/news/breaking-the-code-security-assessment-of-ai-code-agents-through-systematic.jsonld"}}