cd /news/ai-safety/breaking-the-code-security-assessmen… · home › topics › ai-safety › article
[ARTICLE · art-147244] src=aclanthology.org ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks

A TACL 2026 paper by Shoumik Saha, Jifan Chen, Sam Mayers, Sanjay Krishna Gouda, Zijian Wang, and Varun Kumar introduces JAWS-BENCH, a benchmark that measures whether code-capable LLM agents actually compile and run malicious programs rather than merely refusing harmful text. Across seven LLM backends from five families, prompt-only attacks in the empty-workspace regime (JAWS-0) achieved 61% compliance, with 58% harmful, 52% parsing, and 27% running end-to-end; single-file (JAWS-1) compliance reached 100% for stronger models at a mean attack success rate of about 71%, and multi-file (JAWS-M) raised mean ASR to about 75% with 32% runnable attack code. Wrapping an LLM in an agent increased ASR by 1.6x by overturning initial refusals during planning and tool use, with similar trends in SWE-Agent and OpenAI Codex, motivating execution-aware defenses and refusal-preserving agent designs.

read2 min views1 publishedOct 7, 2026
Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
Image: Aclanthology (auto-discovered)
Abstract

Code-capable large language model (LLM) agents are embedded in software engineering workflows where they can read, write, and execute code, raising “jailbreak" stakes beyond text-only settings. Prior evaluations emphasize refusal or harmful-text detection, leaving open whether agents compile and run malicious programs. We present JAWS-BENCH(Jailbreaks Across WorkSpaces), a benchmark spanning three escalating workspace regimes mirroring attacker capability: empty (JAWS-0), single-file (JAWS-1), and multi-file (JAWS-M). We pair it with a hierarchical, executable-aware Judge Framework that tests (i) compliance, (ii) attack success, (iii) syntactic correctness, and (iv) runtime executability to measure de-ployable harm. Across seven LLM backends from five families, prompt-only attacks in JAWS-0 achieve 61% compliance; 58% are harmful, 52% parse, and 27% run end-to-end. In JAWS-1, compliance reaches 100% for stronger models with a mean ASR (Attack Success Rate) ≈ 71%; JAWS-M raises mean ASR to ≈ 75%, with 32% runnable attack code. Wrapping an LLM in an agent increases ASR by 1.6×, by overturning initial refusals during planning and tool use. Additional evaluations with SWE-Agent and OpenAI Codex exhibit similar trends, indicating that JAWS-BENCH can be reused across multiple agent frameworks. Category analyses identify which attack classes are most vulnerable and deployable, motivating execution-aware defenses and refusal-preserving agent designs.

- Anthology ID:
- 2026.tacl-1.94
- Volume:
- [Transactions of the Association for Computational Linguistics, Volume 14](https://aclanthology.org/volumes/2026.tacl-1/)
- Month:
- Year:
  • 2026
  • Address:
  • Cambridge, MA
- Venue:
- [TACL](https://aclanthology.org/venues/tacl/)
- SIG:
- Publisher:
  • MIT Press
- Note:
- Pages:
  • 2081–2102
- Language:
- URL:
- [https://aclanthology.org/2026.tacl-1.94/](https://aclanthology.org/2026.tacl-1.94/)
- DOI:
- [10.1162/tacl.a.792](https://doi.org/10.1162/tacl.a.792)
- Cite (ACL):
- Cite (Informal):
- [Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks](https://aclanthology.org/2026.tacl-1.94/) (Saha et al., TACL 2026)
- PDF:
- [https://aclanthology.org/2026.tacl-1.94.pdf](https://aclanthology.org/2026.tacl-1.94.pdf)
── more in #ai-safety 4 stories · sorted by recency
── more on @jaws-bench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/breaking-the-code-se…] indexed:0 read:2min 2026-10-07 · —