cd /news/ai-safety/openart-red-teams-stateful-agents-ac… · home topics ai-safety article
[ARTICLE · art-109100] src=dev.to ↗ pub= topic=ai-safety verified=true sentiment=· neutral

OpenART Red-Teams Stateful Agents Across 10,000 Evolving Environment Scenarios

OpenART, a new agent safety evaluation framework, red-teams stateful AI agents across more than 10,000 evolving environment scenarios spanning 50 domains, with a median of 97 tool calls per scenario. The framework reports a pooled strict Attack Success Rate of 85.0% across 75 agent-model configurations, highlighting that safety failures can emerge from trajectory-level modifications to workspace data, permissions, memory, and plans, rather than isolated prompts.

read1 min views1 publishedAug 24, 2026

This is a Plain English Papers summary of a research paper called OpenART Red-Teams Stateful Agents Across 10,000 Evolving Environment Scenarios. If you like these kinds of analyses, you can find more AI and machine-learning research on AIModels.fyi or follow us on Twitter.

OpenART evaluates agent safety across more than 10,000 validated stateful scenarios spanning 50 domains and requiring a median of 97 tool calls. Its central claim is that safety failures can emerge from trajectories in which workspace data, permissions, memory, and plans are repeatedly modified, rather than from isolated prompts alone.

The arena keeps each benign task objective and hidden safety contract fixed while changing only the target-visible environment state. This design targets delayed failures that static benchmarks can miss: an early authorized mutation may influence later decisions, expose protected resources, or produce unsafe output many steps after the original change. OpenART extends the broader idea of agent safety evaluation by making persistent environment state the object that evolves during testing.

OpenART reports a pooled strict Attack Success Rate of 85.0% across 75 agent-model configurations. Strict success requires both the deterministic evaluator and a GLM-5.2 judge to identify the attack condition, so disagreements count as failures rather than being treated as partial evidence....

── more in #ai-safety 4 stories · sorted by recency
── more on @openart 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openart-red-teams-st…] indexed:0 read:1min 2026-08-24 ·