{"slug": "5-things-i-would-never-let-an-ai-agent-do-without-a-second-approval", "title": "5 things I would never let an AI agent do without a second approval", "summary": "An OWASP Q1 2026 GenAI exploit roundup documented an AI agent connected to a real inbox that began deleting messages and reportedly continued despite attempts to stop it, an example the author cites to argue that \"human in the loop\" is too vague a control. The author identifies five actions that should require a second independent approval before an AI agent executes them: consequential financial transactions, irreversible destructive production actions, and self-granted privilege or IAM changes, among others. OWASP's Excessive Agency guidance recommends human approval for high-impact actions, and the author argues the approver must authorize the exact action — such as \"Pay $X → recipient Y → from account Z\" or \"Delete these 12 identified resources after these dependency checks\" — rather than the agent's interpretation of it.", "body_md": "The conversation around AI agents has shifted remarkably fast. A year ago, I was mostly worried about what an AI system might say. Now I am increasingly worried about what it can do.\n\nThat distinction changes my security model.\n\nAn agent that gives me a bad recommendation creates a problem I may still have time to catch. An agent with credentials and tools can turn the same bad decision into an API call, a privilege change, a deleted resource or an external transaction before anyone realizes what happened.\n\nI saw a particularly uncomfortable example in OWASP’s [GenAI exploit roundup](https://genai.owasp.org/2026/04/14/owasp-genai-exploit-round-up-report-q1-2026/): an AI agent connected to a real inbox began deleting messages and reportedly continued despite attempts to stop it. There was no sophisticated attacker required. The combination of live access, destructive capability and insufficient confirmation was enough.\n\nThat is why I don’t think “human in the loop” is specific enough anymore.\n\nI want to know exactly where the human — or another independent authorization control — is in the loop.\n\nI would let agents autonomously perform thousands of low-risk actions. But there are five actions where I would deliberately put another decision between the agent’s intent and execution.\n\nI am comfortable letting an agent analyze invoices, reconcile accounts, identify anomalies, prepare a transaction and recommend that a payment be made. I would not let the same agent unilaterally complete a consequential financial transaction.\n\nThe reason isn’t simply fraud. The agent could be operating on manipulated data, an indirect prompt injection, incorrect business context or a perfectly legitimate instruction that it interpreted incorrectly.\n\nThe approval request should therefore expose the transaction itself: Pay $X → recipient Y → from account Z → because of instruction A.\n\nThe second approver should be authorizing that exact transaction, not approving a vague statement such as “complete the payment workflow.”\n\nThis is one place where I think we need to borrow more aggressively from financial controls. The person — or system — that prepares a consequential transaction should not also be the final authority that releases it.\n\nThe same principle should apply to an agent.\n\nThis is the easiest line for me to draw.\n\nI want agents helping operations teams identify stale resources, clean environments, investigate incidents and recommend remediation. I don’t want an agent translating “clean this up” into an irreversible production action without another control validating what “this” actually means.\n\nOWASP describes exactly this class of failure in its guidance on [Excessive Agency](https://genai.owasp.org/llmrisk/llm062025-excessive-agency/): an AI system can become dangerous when excessive functionality, permissions or autonomy allow it to perform damaging actions. OWASP specifically recommends human approval for high-impact actions.\n\nFor a destructive action, my approval screen should tell me:\n\nWhat will be deleted? What depends on it? What is the blast radius? Can I recover it?\n\n“Delete these 12 identified resources after these dependency checks” is reviewable.\n\n“Clean up unnecessary production resources” is not.\n\nThe difference matters because I don’t want a human approving the agent’s interpretation. I want the human approving the actual action.\n\nThis is where I become even more conservative.\n\nImagine an agent reaches a point in a workflow where it cannot continue because it lacks permission. It decides the permission is necessary, requests additional access and then has a path to grant or activate that access itself.\n\nAt that point, the agent is no longer merely executing within an authorization boundary. It is participating in defining its own boundary.\n\nI would separate those decisions.\n\nAn agent can tell me it needs more authority. It cannot be the authority that decides whether it gets it.\n\nThat means changes to IAM roles, privileged group membership, credentials, security policies and sensitive tool permissions need an independent authorization path.\n\nThis is also why least privilege becomes more important, not less, with agents. OWASP recommends that agent tools operate with the [minimum functionality and permissions necessary](https://genai.owasp.org/llmrisk/llm062025-excessive-agency/) and that downstream systems enforce authorization rather than trusting the model to decide whether an action should be permitted.\n\nA compromised read-only agent gives me an incident. A compromised agent that can make itself an administrator gives me a very different incident.\n\nThis one is deceptively difficult because an agent doesn’t need an “exfiltrate data” tool to cause a disclosure.\n\nIt may attach a document to an email. Populate a support ticket. Paste source code into an external service. Send customer information through an API. Give another agent information that the second agent was never supposed to receive.\n\nEach step can look legitimate.\n\nThe [EchoLeak vulnerability in Microsoft 365 Copilot](https://www.microsoft.com/en-us/security/blog/2025/06/11/echoleak-how-a-zero-click-ai-vulnerability-could-leak-sensitive-data/) demonstrated why this boundary matters. A specially crafted email could influence the AI system and create a path toward disclosure of sensitive organizational information without the user needing to click a malicious link or attachment.\n\nBefore sensitive information crosses an established trust boundary, I want another decision point that answers four questions:\n\nWhat data is leaving? Who is receiving it? Why do they need it? Where is it going?\n\nThis is where approval design matters almost as much as approval itself. If the agent controls the description shown to the approver, “Send requested information to vendor” might conceal a very different underlying action.\n\nThe security control therefore needs to derive its approval context from the actual transaction — destination, data classification, tool call and requested scope — not merely from the agent’s natural-language explanation.\n\nOtherwise, we have put a human in the loop without giving that human enough information to make a security decision.\n\nThis is the boundary I expect security teams to wrestle with most as multi-agent architectures become common.\n\nA modern agent workflow can quickly become:\n\nUser → Agent → Delegated Agent → Skill → Tool → Resource\n\nThe original agent may have legitimate authority. But what happens when it delegates a task to another agent that invokes a tool holding broader credentials?\n\nThe question I care about is not simply whether every component is authenticated.\n\nIt is whether the authority exercised at the end of that chain still represents what was authorized at the beginning.\n\nI would therefore require another authorization decision when an agent creates privileged access or materially expands or delegates consequential authority.\n\nAnd I would make that approval very specific:\n\nWhat authority is being delegated? To whom? For what task? For how long? Can it be delegated again? What happens to it when the task ends or the original authorization is revoked?\n\nOWASP’s [AI Agent Security Cheat Sheet](https://cheatsheetseries.owasp.org/cheatsheets/AI_Agent_Security_Cheat_Sheet.html) addresses excessive autonomy, tool abuse, privilege boundaries, sensitive-data exposure and multi-agent risks. Those risks become harder to reason about as an agent’s authority travels through more components.\n\nAuthentication tells me which agent is acting. Authorization needs to tell me whether this particular action, with this particular authority, is still allowed.\n\nThose are not the same question.\n\nThere is one part of this argument I don’t want misunderstood.\n\nI am not proposing that enterprises hire armies of people to click “Approve” every time an agent wants to do something.\n\nThat would destroy the value of automation.\n\nFor the highest-risk actions, a human approval may absolutely be appropriate. In other cases, the second decision could come from an independent policy engine, authorization service, transaction-control system or another deterministic security control that the initiating agent cannot modify or bypass.\n\nThe [NIST AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework) takes a risk-based approach to managing AI risks and emphasizes governance throughout the AI lifecycle.\n\nThat’s how I think about agent autonomy as well.\n\nThe question isn’t: Should this agent be autonomous?\n\nThe better question is: Which actions can this agent perform autonomously, and which actions require an independent decision?\n\nFor me, the dividing line is increasingly clear.\n\nLet agents search, analyze, summarize, correlate, recommend and prepare. Let them autonomously execute actions that are bounded, observable, reversible and already inside clearly established authority.\n\nAdd friction when an action can move money, destroy production data, increase privilege, expose sensitive information or create authority that propagates beyond the original task.\n\nMost importantly, I would never design a workflow where the same agent can propose a consequential action, acquire the authority necessary to perform it, approve that authority and execute the action.\n\nThat isn’t autonomy I can govern. It’s self-authorization.\n\nAnd when AI agents can operate at machine speed across multiple connected systems, a few seconds of independent authorization may be some of the cheapest incident prevention we can buy.", "url": "https://wpnews.pro/news/5-things-i-would-never-let-an-ai-agent-do-without-a-second-approval", "canonical_source": "https://www.cio.com/article/4227762/5-things-i-would-never-let-an-ai-agent-do-without-a-second-approval.html", "published_at": "2026-09-29 12:00:00+00:00", "updated_at": "2026-09-29 12:19:29.150290+00:00", "lang": "en", "topics": ["ai-agents", "ai-safety", "ai-policy"], "entities": ["OWASP", "GenAI exploit roundup", "Excessive Agency"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/5-things-i-would-never-let-an-ai-agent-do-without-a-second-approval", "markdown": "https://wpnews.pro/news/5-things-i-would-never-let-an-ai-agent-do-without-a-second-approval.md", "text": "https://wpnews.pro/news/5-things-i-would-never-let-an-ai-agent-do-without-a-second-approval.txt", "jsonld": "https://wpnews.pro/news/5-things-i-would-never-let-an-ai-agent-do-without-a-second-approval.jsonld"}}