cd /news/ai-agents/why-approval-prompts-don-t-work-as-a… · home › topics › ai-agents › article
[ARTICLE · art-148320] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Why approval prompts don't work as a security boundary for coding agents

A developer building Cirvix AgentControl, an open-source default-deny policy layer for agent tool calls, argues that human approval prompts are not a real security boundary for coding agents because approvals suffer from fatigue, never expire, and can be replayed against drifted environments. The proposed fix binds each approval to a SHA-256 hash of the exact action plus a short expiry window, so the enforcement layer verifies a one-time ticket at execution rather than a standing credential.

by read3 min views2 publishedOct 9, 2026

When a coding agent asks a human to approve a file change, a database call, or a deploy, the approval prompt feels like a security boundary. It's not.

In practice, approval prompts leak authority for three reasons: approval fatigue, approvals that never expire, and stale pending approvals that outlive the decision they capture.

An agent that runs an entire PR lifecycle can surface dozens of prompts in a single session. Most of them are low-risk: formatting a config, reading a log, re-running tests. A few are high-risk: writing to a deployment manifest, touching secrets, changing network rules.

When everything looks the same in the prompt, humans stop reading. They learn to click "Approve" to keep the agent moving and catch up later. That later is when the deploy went to the wrong environment and the rollback took thirty minutes.

The fix isn't more prompts. It's making every prompt distinguishable.

Most agent toolkits record an approval as a boolean decision on a category of action: "approve kubectl apply" or "approve writes to /etc". There is no time bound attached.

An approval granted at 9:00 AM should not authorize the same action at 4:00 PM, after the incident context has changed, the on-call has handed off, and the agent's mission scope has drifted.

Without an expiry, the approval becomes a standing credential. The human reviewed a point-in-time description and unknowingly minted a persistent grant.

The flip side: a prompt sits in a queue while the human is on a call or asleep. Three hours later, the agent replays the pending request against a newer snapshot of the environment.

The description the human approved no longer matches what would execute. The approval was sound when issued; it is unsound when fulfilled.

This is why "pending approvals" and "authorized actions" are not the same thing. A pending approval captures a decision about a specific action at a specific time. Once that time window closes, the decision must be withdrawn.

A practical fix that keeps humans in the loop without minting standing credentials:

This is how Cirvix AgentControl structures an approval:

import { createHash } from "crypto";

function actionHash(action: Record<string, unknown>): string {
  return createHash("sha256")
    .update(JSON.stringify(action, null, 0))
    .digest("hex");
}

interface HoldApproval {
  id: string;
  actionHash: string;
  expiresAt: number; // epoch ms
  approvedBy: string;
  approvedAt: number;
  status: "pending" | "approved" | "denied" | "expired";
}

function verifyApproval(
  hold: HoldApproval,
  requestedAction: Record<string, unknown>
): boolean {
  const freshHash = actionHash(requestedAction);

  return (
    hold.status === "approved" &&
    hold.actionHash === freshHash &&
    Date.now() < hold.expiresAt
  );
}

const call = { tool: "kubectl_apply", manifest: "deploy.yaml", namespace: "prod" };
const hold: HoldApproval = {
  id: "hold-001",
  actionHash: actionHash(call),
  expiresAt: Date.now() + 10 * 60 * 1000,
  approvedBy: "human@example.com",
  approvedAt: Date.now(),
  status: "approved",
};

console.log(verifyApproval(hold, call)); // true

const drifted = { ...call, namespace: "prod-backup" };
console.log(verifyApproval(hold, drifted)); // false

The three checks — status approved, hash matches the exact action, and current time is still inside the expiry — make the approval a verifiable ticket for one execution, not a standing credential.

Humans still review. The difference is that their review is scoped: a short window, a specific action, and a clear failure mode when the scope shifts.

Prompt-based approvals are still useful as a UX. They just shouldn't be the boundary. The boundary is the hash-bound, time-bound decision that the enforcement layer verifies at execution.

I'm building Cirvix AgentControl, an open-source default-deny policy layer for agent tool calls: https://github.com/CIRVIX/agent-control (try npx @cirvix_ai/agent-control scan).

── more in #ai-agents 4 stories · sorted by recency
── more on @cirvix agentcontrol 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/why-approval-prompts…] indexed:0 read:3min 2026-10-09 · —