arXiv:2610.00797v1 Announce Type: new Abstract: Contextual security defenses prevent AI agents from taking rogue actions by synthesizing a task-specific policy and enforcing it on the agent's tool calls. In multi-step tasks, however, which actions are valid often depends on what the agent has already done and learned. We present Sapien, a policy engine for enforcing stateful contextual policies. A Sapien policy specifies permitted tool-call sequences using a regular expression extended with stateful predicates, deferred policy generation, and scoped semantic checks. We show that Sapien stays within a few percent of an unconstrained agent's utility. Even if the agent is fully hijacked, Sapien's policies rule out 93-95% of attacks on AgentDojo and 62-85% on Toolathlon (twice as many as tool allowlists on long-horizon tasks).
Sapien: A Stateful Policy Engine for Autonomous AI Agents
Researchers introduced Sapien, a stateful policy engine that enforces contextual security policies on autonomous AI agents' tool calls by specifying permitted tool-call sequences with regular expressions extended with stateful predicates, deferred policy generation, and scoped semantic checks. According to the arXiv paper 2610.00797v1, Sapien stays within a few percent of an unconstrained agent's utility while ruling out 93-95% of attacks on AgentDojo and 62-85% on Toolathlon even when the agent is fully hijacked, roughly twice as many as tool allowlists on long-horizon tasks.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.