{"slug": "new-archestra-s-openappa-saturates-two-major-security-benchmarks-with-a-0-attack", "title": "New Archestra's OpenAPPA Saturates Two Major Security Benchmarks with a 0% Attack Success Rate", "summary": "Archestra released OpenAPPA, an open-source security engine that runs outside an AI agent's prompt and execution loop, reporting a 0% attack success rate on the Bench-Corp (20 multi-step enterprise workflows) and AgentThreatBench security benchmarks, versus 10% for Claude Code's auto mode and 31% for Microsoft FIDES. OpenAPPA implements an Agentic Permissions Policy Algebra (APPA), described in an arXiv paper by Arseny Kravchenko, Vadim Liventsev, Innokentii Konstantinov, Ildar Iskhakov, and Matvey Kukuy, using a single appa.toml configuration file that defines data sources, audiences, trust levels, and authorities with deterministic enforcement rules. Archestra argues stochastic judge models cannot track data flow across tool calls and top out at 99.3% detection, leaving 0.7% of millions of calls as breaches.", "body_md": "Archestra released [OpenAPPA](https://www.openappa.com/), an open-source security engine designed to stop data exfiltration caused by prompt injection or model hallucination. OpenAPPA runs outside the agent’s prompt and execution loop. Its configuration details concepts such as data sources, [audiences](https://www.openappa.com/contracts#audiences), [trust levels](https://www.openappa.com/contracts#trust), and [authorities](https://www.openappa.com/contracts#authorities), along with their associated deterministic security enforcement rules. The team reports zero successful attacks when running security benchmarks [Bench-Corp](https://github.com/archestra-ai/OpenAPPA/tree/main/bench/corp) (20 multi-step enterprise workflows) and [AgentThreatBench](https://github.com/UKGovernmentBEIS/inspect_evals/tree/main/src/inspect_evals/agent_threat_bench), vs. 10% for [Claude Code’s auto mode](https://code.claude.com/docs/en/auto-mode-config) and 31% for [Microsoft FIDES](https://learn.microsoft.com/en-us/agent-framework/agents/security).\n\nOpenAPPA’s documentation explains why a stochastic approach to automated policy enforcement fails:\n\nThe industry’s answer to approval fatigue is a second model that judges each tool call: Claude Code’s auto mode, Codex’s auto-review, and other [auto-modes](https://www.openappa.com/openappa-vs-auto-mode).\n\nBy design, they cannot track data flow across tool calls. Because classifiers are prompt-injectable themselves, harnesses hide tool outputs from them, so the judge never sees the data at all.\n\nBecause of their probabilistic design, even the best top out at [99.3%](https://openai.github.io/openai-guardrails-python/ref/checks/prompt_injection_detection/): at millions of calls, 0.7% is a lot of breaches.\n\n[…] Rule sets end up either so tight they break the agent or so intricate nobody can audit what they permit.\n\nOn the one hand, agents have proven skilled at working around simple but common approaches like allowlists or denylists of tools: a denied `rm -rf` may be replaced by an equivalent Python script. On the other hand, extending denylists or overly restricting policies to protect against eager agents results in decreased utility (e.g., while the agent does not leak data, it does not perform the task successfully because of the restrictions). OpenAPPA’s GitHub repository reminds developers:\n\nAgent security has two axes: an agent that permits unauthorized flows is unsafe, and an agent that refuses valid work is useless.\n\nArchestra seeks to resolve the tension between strict enforcement and operational utility with what it calls an Agentic Permissions Policy Algebra (APPA), [described in a paper](https://arxiv.org/abs/2607.24625) by [Arseny Kravchenko](https://arxiv.org/search/cs?searchtype=author&query=Kravchenko,+A), [Vadim Liventsev](https://arxiv.org/search/cs?searchtype=author&query=Liventsev,+V), [Innokentii Konstantinov](https://arxiv.org/search/cs?searchtype=author&query=Konstantinov,+I), [Ildar Iskhakov](https://arxiv.org/search/cs?searchtype=author&query=Iskhakov,+I), and [Matvey Kukuy](https://arxiv.org/search/cs?searchtype=author&query=Kukuy,+M).\n\nOpenAPPA implements this approach with a pluggable engine that is executed outside the agent’s loop, thus defeating any attempts by the underlying language model to inspect, negotiate with, or bypass policy rules. The security policies take the form of a single [`appa.toml`](https://www.openappa.com/policy-configuration) configuration file that details data sources, [audiences](https://www.openappa.com/contracts#audiences), [trust levels](https://www.openappa.com/contracts#trust), and [authorities](https://www.openappa.com/contracts#authorities).\n\nOpenAPPA jointly labels and monitors both audience (the authorized set of consumers) and trust (the degree of data verification). Labels compose monotonically using lattice algebra; labels can only become more restrictive; reading restricted records narrows the audience; reading unvetted external web pages lowers trust.\n\nEach tool contract defines three primary operational attributes: `requires` (the audience membership and trust levels necessary to run the tool), `delta` (the restrictions applied when the tool returns data), and `effects` (an audit trail of successful actions).\n\nIn the following example, reading a ticket via `get_ticket_from_crm` restricts the trajectory’s audience to `internal`. The subsequent `process_internal_data` call requires the audience to be within `internal`.\n\n```\n[[policy.tool]]\nname = \"get_ticket_from_crm\"\ndelta = { audience = [\"internal\"] }\n\n[[policy.tool]]\nname = \"publish_update\"\nrequires = { audience = { contains = [\"public\"] } }\n\n[[policy.tool]]\nname = \"process_internal_data\"\nrequires = { audience = { within = [\"internal\"] } }\n```\n\nIn the following example, the `read_web_page` tool’s result is marked as `suspicious`. Upon reception of a result from the `read_web_page` tool, OpenAPPA blocks `apply_db_migration` because it requires `trusted` data.\n\n```\n[policy]\nversion = 2\ntrust_chain = [\"suspicious\", \"trusted\"]\n\n[[policy.tool]]\nname = \"read_web_page\"\ndelta = { trust = \"suspicious\" }\n\n[[policy.tool]]\nname = \"apply_db_migration\"\nrequires = { trust = \"trusted\" }\n```\n\nOpenAPPA additionally has explicit recovery semantics. When an agent attempts an illegal action, the engine halts dispatch and provides structured pathways to proceed. *Sanitizers* may edit payloads, e.g., stripping personally identifiable information, to safely expand the permitted audience. *Authorities* route requests to human operators or internal verification APIs for single-action approval. *Disposable Child Branches* enable on-demand confinement: when an agent must ingest untrusted data, the engine isolates the read in a transient subagent branch, returning only schema-attested, sanitized outputs to the parent.\n\nThe [Bench-Corp and OWASP AgentThreatBench](https://www.openappa.com/evaluation) benchmarks show that OpenAPPA maintains high utility while under strict security constraints, reporting a 0% attack success rate and an 89% task completion rate. Claude Code’s native auto mode yielded a 10% attack success rate with a 90% completion rate. Microsoft FIDES permitted 31% of attacks to succeed and completed only 41% of tasks. The team of researchers reports in the paper that ablation experiments seem to validate the value of recovery strategies: task completion fell to 35.0% when remedy plans were completely disabled.\n\nBench-Corp and AgentThreatBench test [explicit policy breaches](https://www.openappa.com/how-it-works#the-core-concepts): sensitive data sharing, prompt injection, approval and ordering, and tenant isolation. Bench-Corp is a highly specialized corporate-assistant benchmark designed to evaluate how security policies hold up across complex, multi-step enterprise workflows. [AgentThreatBench](https://dev.to/vaishnavi_gudur/agentthreatbench-the-first-owasp-agentic-top-10-security-benchmark-6pp) operationalizes the [OWASP Top 10 for Agentic Applications (2026)](https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/) into executable tasks. It was recently merged into the official [UK AI Safety Institute’s `inspect_evals` repository](https://github.com/UKGovernmentBEIS/inspect_evals/pull/1037). AgentThreatBench uses a dual-metric scoring system, scoring both utility and security. OpenAPPA achieves a perfect security score on the benchmark with zero successful attacks.\n\nOpenAPPA is currently a preview. The formal algebra and recovery guarantees are published on [arXiv (APPA: Recoverable Information-Flow Control for Real-World LLM Agents)](https://arxiv.org/abs/2607.24625). The reader is encouraged to read the paper and the documentation, both of which contain plenty of additional technical details, accompanying illustrations, and empirical results.", "url": "https://wpnews.pro/news/new-archestra-s-openappa-saturates-two-major-security-benchmarks-with-a-0-attack", "canonical_source": "https://www.infoq.com/news/2026/10/open-APPA-zero-security-breach/?utm_campaign=infoq_content&utm_source=infoq&utm_medium=feed&utm_term=global", "published_at": "2026-10-03 23:41:00+00:00", "updated_at": "2026-10-04 00:06:50.874465+00:00", "lang": "en", "topics": ["ai-safety", "ai-agents", "ai-tools", "artificial-intelligence", "ai-research"], "entities": ["Archestra", "OpenAPPA", "Agentic Permissions Policy Algebra", "Claude Code", "Microsoft FIDES", "Bench-Corp", "AgentThreatBench", "Arseny Kravchenko"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/new-archestra-s-openappa-saturates-two-major-security-benchmarks-with-a-0-attack", "markdown": "https://wpnews.pro/news/new-archestra-s-openappa-saturates-two-major-security-benchmarks-with-a-0-attack.md", "text": "https://wpnews.pro/news/new-archestra-s-openappa-saturates-two-major-security-benchmarks-with-a-0-attack.txt", "jsonld": "https://wpnews.pro/news/new-archestra-s-openappa-saturates-two-major-security-benchmarks-with-a-0-attack.jsonld"}}