{"slug": "why-approval-prompts-don-t-work-as-a-security-boundary-for-coding-agents", "title": "Why approval prompts don't work as a security boundary for coding agents", "summary": "A developer building Cirvix AgentControl, an open-source default-deny policy layer for agent tool calls, argues that human approval prompts are not a real security boundary for coding agents because approvals suffer from fatigue, never expire, and can be replayed against drifted environments. The proposed fix binds each approval to a SHA-256 hash of the exact action plus a short expiry window, so the enforcement layer verifies a one-time ticket at execution rather than a standing credential.", "body_md": "When a coding agent asks a human to approve a file change, a database call, or a deploy, the approval prompt feels like a security boundary. It's not.\n\nIn practice, approval prompts leak authority for three reasons: approval fatigue, approvals that never expire, and stale pending approvals that outlive the decision they capture.\n\nAn agent that runs an entire PR lifecycle can surface dozens of prompts in a single session. Most of them are low-risk: formatting a config, reading a log, re-running tests. A few are high-risk: writing to a deployment manifest, touching secrets, changing network rules.\n\nWhen everything looks the same in the prompt, humans stop reading. They learn to click \"Approve\" to keep the agent moving and catch up later. That later is when the deploy went to the wrong environment and the rollback took thirty minutes.\n\nThe fix isn't more prompts. It's making every prompt distinguishable.\n\nMost agent toolkits record an approval as a boolean decision on a category of action: \"approve kubectl apply\" or \"approve writes to /etc\". There is no time bound attached.\n\nAn approval granted at 9:00 AM should not authorize the same action at 4:00 PM, after the incident context has changed, the on-call has handed off, and the agent's mission scope has drifted.\n\nWithout an expiry, the approval becomes a standing credential. The human reviewed a point-in-time description and unknowingly minted a persistent grant.\n\nThe flip side: a prompt sits in a queue while the human is on a call or asleep. Three hours later, the agent replays the pending request against a newer snapshot of the environment.\n\nThe description the human approved no longer matches what would execute. The approval was sound when issued; it is unsound when fulfilled.\n\nThis is why \"pending approvals\" and \"authorized actions\" are not the same thing. A pending approval captures a decision about a specific action at a specific time. Once that time window closes, the decision must be withdrawn.\n\nA practical fix that keeps humans in the loop without minting standing credentials:\n\nThis is how Cirvix AgentControl structures an approval:\n\n``` js\nimport { createHash } from \"crypto\";\n\nfunction actionHash(action: Record<string, unknown>): string {\n  return createHash(\"sha256\")\n    .update(JSON.stringify(action, null, 0))\n    .digest(\"hex\");\n}\n\ninterface HoldApproval {\n  id: string;\n  actionHash: string;\n  expiresAt: number; // epoch ms\n  approvedBy: string;\n  approvedAt: number;\n  status: \"pending\" | \"approved\" | \"denied\" | \"expired\";\n}\n\nfunction verifyApproval(\n  hold: HoldApproval,\n  requestedAction: Record<string, unknown>\n): boolean {\n  const freshHash = actionHash(requestedAction);\n\n  return (\n    hold.status === \"approved\" &&\n    hold.actionHash === freshHash &&\n    Date.now() < hold.expiresAt\n  );\n}\n\nconst call = { tool: \"kubectl_apply\", manifest: \"deploy.yaml\", namespace: \"prod\" };\nconst hold: HoldApproval = {\n  id: \"hold-001\",\n  actionHash: actionHash(call),\n  expiresAt: Date.now() + 10 * 60 * 1000,\n  approvedBy: \"human@example.com\",\n  approvedAt: Date.now(),\n  status: \"approved\",\n};\n\nconsole.log(verifyApproval(hold, call)); // true\n\nconst drifted = { ...call, namespace: \"prod-backup\" };\nconsole.log(verifyApproval(hold, drifted)); // false\n```\n\nThe three checks — status approved, hash matches the exact action, and current time is still inside the expiry — make the approval a verifiable ticket for one execution, not a standing credential.\n\nHumans still review. The difference is that their review is scoped: a short window, a specific action, and a clear failure mode when the scope shifts.\n\nPrompt-based approvals are still useful as a UX. They just shouldn't be the boundary. The boundary is the hash-bound, time-bound decision that the enforcement layer verifies at execution.\n\nI'm building Cirvix AgentControl, an open-source default-deny policy layer for agent tool calls: [https://github.com/CIRVIX/agent-control](https://github.com/CIRVIX/agent-control) (try `npx @cirvix_ai/agent-control scan`).", "url": "https://wpnews.pro/news/why-approval-prompts-don-t-work-as-a-security-boundary-for-coding-agents", "canonical_source": "https://dev.to/umangcirvix/why-approval-prompts-dont-work-as-a-security-boundary-for-coding-agents-4h8e", "published_at": "2026-10-09 14:40:53+00:00", "updated_at": "2026-10-09 14:51:44.062075+00:00", "lang": "en", "topics": ["ai-agents", "ai-safety", "ai-tools", "developer-tools"], "entities": ["Cirvix AgentControl", "GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/why-approval-prompts-don-t-work-as-a-security-boundary-for-coding-agents", "markdown": "https://wpnews.pro/news/why-approval-prompts-don-t-work-as-a-security-boundary-for-coding-agents.md", "text": "https://wpnews.pro/news/why-approval-prompts-don-t-work-as-a-security-boundary-for-coding-agents.txt", "jsonld": "https://wpnews.pro/news/why-approval-prompts-don-t-work-as-a-security-boundary-for-coding-agents.jsonld"}}