cd /news/ai-safety/execution-grounded-security-testing-… · home topics ai-safety article
[ARTICLE · art-76390] src=arxiv.org ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Execution-Grounded Security Testing for Coding Agents in Software Engineering Pipelines

A new execution-grounded red-team testing framework reveals that coding agents integrated into system operations can be induced to carry out unsafe actions on the surrounding environment, with verified unsafe execution reaching 73.61% on code carriers and 53.93% on text carriers. The framework, presented in arXiv:2607.22569v1, embeds target unsafe operations into routine software engineering workloads and uses observable sandbox evidence to probe the execution-layer security boundary. The results indicate that coding agents in system operations remain insecure under task disguise and demand stronger security testing and safeguards.

read1 min views1 publishedJul 28, 2026

arXiv:2607.22569v1 Announce Type: new Abstract: Coding agents are increasingly integrated into system operations, where their tool use can directly modify project artifacts, execution environments, and the underlying system. For example, if a coding agent inserts a hook into a system startup or configuration script, that change can persist after the interaction, be triggered later, and abuse delegated user or system privileges to modify the system. This makes security testing a system problem: the key question is not only what the agent says, but what it actually does to the surrounding environment. We present an execution-grounded red-team testing framework for probing this execution-layer security boundary using observable sandbox evidence, including tool invocations, runtime traces, and file-system diffs. Our framework embeds target unsafe operations into routine software engineering workloads, including unit testing, regression testing, crash reproduction, and validation, and uses an execution oracle to guide refinement when an initial probe is rejected or fails. Across multiple agent frameworks and model backbones, our red-team workload reformulation substantially increases verified unsafe execution, reaching 73.61% on code carriers and 53.93% on text carriers. These results show that coding agents in system operations remain insecure under task disguise: once risky intent is hidden inside plausible engineering tasks, the agent can be induced to carry out unsafe actions on the surrounding system. More broadly, coding agents in system operations still demand stronger security testing and safeguards.

── more in #ai-safety 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/execution-grounded-s…] indexed:0 read:1min 2026-07-28 ·