That moment does not exist by default. MCP standardizes how agents discover and call tools, but it deliberately leaves approval to the people building the system. The specification says a human should be able to deny a tool call, then stops there. The safety flags a server can attach to a tool are hints, not locks, and a client that ignores them breaks nothing.
Human in the loop approval in MCP is a design choice, not a switch you flip. Learning how to add approval gates to MCP server actions comes down to one principle: the gate has to sit where the side effect happens, enforced by code the agent cannot route around. Get that right and the rest becomes engineering.
Pausing every tool call for a human is not a safety strategy. It trains reviewers to click Approve without reading. The goal is to spend human attention only where a mistake is expensive.
A practical way to decide is to sort every action an agent can take into tiers, using three questions:
Environment matters as much as the action itself. The same delete call that is harmless against a sandbox deserves a gate once it points at live customer data, so production versus staging should be part of the policy.
Most integration layers encode this thinking as a policy per integration. Corsair, for example, labels every endpoint as read, write, or destructive, then maps those labels to allow, require approval, or deny through its permission modes: open, cautious, strict, and readonly. Cautious lets agents read and write freely but sends destructive calls to a human. Strict also gates ordinary writes and blocks destructive actions outright.
Whatever tooling you use, write your tiers down before you write any gating code, because the tiers are the policy.
MCP tool definitions can carry annotations such as readOnlyHint, destructiveHint, idempotentHint, and openWorldHint. They look like a ready-made risk system, and it is tempting to wire them straight into approval logic. They are not enforcement, for four reasons.
Use annotations for what they are good at, which is helping clients decide what to show a user. Then enforce on the server.
A real approval gate has five properties:
The shape of the code is small:
async function runAction(call, caller) {
const policy = policyFor(call.action); // allow, require_approval, or deny
if (policy === "deny") return blocked("Blocked by policy");
if (policy === "allow") return execute(call);
const approved = await approvals.consume(
caller,
call.action,
hash(call.args)
);
if (!approved) {
await approvals.createPending(caller, call.action, call.args);
return pendingApproval();
}
return execute(approved.frozenArgs); // runs once, with the reviewed arguments
}
Two details matter more than they look.
First, put the gate at the layer that owns the integration, not at a single tool. If a delete can be reached through an MCP tool, a background job, or a script, all three must pass the same check, otherwise gating one route simply teaches the agent to take another.
Second, deny by default for anything unclassified. A new tool with no tier stays gated until someone decides what it is.
This is the reasoning behind Corsair's MCP adapters. They give the agent three tools for listing operations, inspecting schemas, and running scripts, and the permission policy gates the calls that run through them. The check sits at the integration endpoint, so it does not depend on how any one tool describes itself.
Once the gate exists, the real work is the lifecycle around it. Every human in the loop approval flow has three phases.
When a call needs approval, the server does not execute it and does not guess. It records a pending request and tells the agent what happened.
There are two ways to hold the call:
Corsair supports both, with asynchronous as the default, and its hosted approval page gives reviewers an approve or deny screen without building one yourself. The decision record stays in your own database either way.
Give every pending request an expiry and treat silence as a no. Corsair applies a ten-minute window unless you set another, and its docs recommend denying on timeout.
After approval there are two clean ways to continue.
The agent can retry the same call and the server matches it to the approved record, or the server can execute the stored request itself the moment the decision lands.
In both cases, the action runs with the arguments the reviewer saw. Never ask the model to rebuild the call from memory, because a slightly different argument is a different action.
An approval flow without a record is just a delay.
For every gated call, capture:
Keep approval and success as separate events. A yes does not mean the delete worked, because the downstream API can still reject it or time out.
The spec expects clients to log tool use for audit, but a client log is not your log. Record events on the server, using hooks that run before and after each call or equivalent middleware, so no route can skip the record.
The July 28, 2026 revision of the MCP specification changes how a server can ask a person for input in the middle of a tool call, which makes it the most important update for anyone building approval workflows with MCP servers.
Before, a server that wanted a confirmation sent its own request back to the client over a connection that had to stay open. The new revision makes the protocol stateless, drops the initialize handshake and sessions, and replaces those server-initiated requests with Multi Round Trip Requests, usually shortened to MRTR.
Any request can land on any server instance, which suits approval flows that may wait on a human.
The flow works like this:
input_required. The result carries one or more input requests, such as an elicitation asking the user a question, and optionally an opaque requestState.
Here is what an approval request looks like on the wire:
{
"resultType": "input_required",
"inputRequests": {
"approve_delete": {
"method": "elicitation/create",
"params": {
"mode": "form",
"message": "Delete 214 inactive customer records from production?",
"requestedSchema": {
"type": "object",
"properties": {
"approve": {
"type": "boolean"
}
},
"required": ["approve"]
}
}
}
},
"requestState": "signed, expiring blob"
}
The retry carries the person's answer in inputResponses, next to the same requestState.
Elicitation has three possible outcomes: accept, decline, and cancel. Treat only an accept with approve set to true as permission to run the action, and handle the other two as a clean refusal.
Four rules keep this flow safe:
Know the limits of in-band approval.
An elicitation answer travels back through the client, and the spec only says clients should offer approval controls, so a client that automatically accepts can answer yes on a human's behalf. Form mode also exposes the exchange to the client and the model's context, and it must never be used to collect secrets.
For changes to production data, a safer pattern is URL mode elicitation, added in the November 2025 revision. It sends the reviewer to a page on your own domain, where your server authenticates them. The client only learns that the user agreed to open the link, and your server learns who actually approved.
In practice, many teams use both: a quick in-band confirmation for lower tiers, and an out-of-band page for tiers four and five.
Reviews that take hours do not fit a single held request at all. For those, return a pending result and resume later, or look at the Tasks extension, which moved out of the core protocol in this revision and lets clients poll for the status of long-running work.
Older clients may not speak the new revision yet, so keep the asynchronous pattern from the previous section as a fallback.
Put the earlier sections together and the whole policy fits on one page.
Approval fatigue is a security problem in its own right.
When reviewers face a flood of low-stakes prompts, they stop reading them. Show the exact action in plain language, include the target and the count, save prompts for the tiers where a mistake is costly, and let routine safe actions pass with logging instead.
Approval gates work best when they live in the integration layer, close to the API call they protect, instead of being rebuilt inside every MCP server.
Corsair is an open source TypeScript integration layer for AI agents that applies permission modes per integration, holds risky calls for human approval with frozen arguments, and keeps the decision record in your own database. It connects to agents through MCP adapters, so the same policy covers every route an agent can take.
Start with one destructive action, gate it, and test approval, denial, and expiry before you widen the policy.
Classify each action by risk, then enforce a policy in server code right before the side effect. Calls that need approval create a pending record, return a clear message to the agent, and run only after a verified person approves, using the arguments they reviewed. Log the request, the decision, and the result.
destructiveHint enough to gate dangerous actions?
No. The specification treats annotations as hints and tells clients to consider them untrusted unless the server is trusted. Use them to help clients decide what to display, and enforce approval on the server where the action actually executes.
Multi Round Trip Requests, introduced in the July 28, 2026 specification revision, let a server end a tool call with an input_required result and ask the client for input. The client collects the answer and retries the same call with it attached. This lets a server request confirmation without holding a connection open or storing session state.
Use elicitation for quick confirmations on lower-risk actions in interactive clients. For production data and high-value actions, use an out-of-band review page, or URL mode elicitation, where your server authenticates the approver. In-band answers pass through the client, so they are weaker evidence of who actually decided.
Gate only the tiers where mistakes are costly, usually external-facing, destructive, and irreversible actions. Show the exact action, target, and record count in plain language, expire stale requests, and let routine low-risk actions run with logging. If reviewers approve nearly everything, tighten the policy or improve what the prompt shows.