28 payments got through. 12 didn't. The cap was $200 a day, and we fired all 40 at the exact same moment on purpose, because we don't trust our own code more than we'd trust yours.
Disclosure: I work on Pink at PinkWallet. Every number below came from a live run against our public sandbox, reproducible with the script linked at the bottom.
That race is the easy failure mode to catch. The harder one: even a cap that holds doesn't matter if the agent is still sitting on a real payment credential it can use on its own. Pink Agentic AI Payments closes both, by keeping the check and the credential off the agent's side of the fence entirely. Here's what each failure looks like and how it's closed.
A spending cap usually gets enforced like this, somewhere in application code or inside the agent's own tool call: read the amount spent so far, check it against the limit, and if it's under, let the payment go through. That works for one call at a time. It falls apart the moment an agent (or a swarm of agents, or retried tool calls) fires several payment requests in parallel.
Here's why. Say the limit is $200/day and $190 has already been spent. Ten parallel calls each ask for $7. Each one reads "spent so far: $190," each one computes $190 + $7 = $197, which is under $200, so each one passes its own check. All ten go through. The agent just spent $70 against a $10 remaining budget, and nothing in that check-then-spend logic ever saw it coming, because the check and the spend were not the same atomic step.
This isn't hypothetical. It's the standard "check-then-act" race condition, and it gets worse, not better, as agents get faster and start running tool calls concurrently to save latency.
Even if you fix the race, a second problem remains: where does the actual payment credential live? In most setups, the agent process holds a real API key, card number, or wallet credential, and the "spending limit" is a piece of logic that runs in the same process, checking the same memory, before the agent is allowed to use that credential.
The issue is that anything enforced inside the agent's own context can potentially be bypassed by what's inside that context, including text the agent reads from a tool result, a webpage, or a document, if that text is crafted to look like an instruction. If the limit-check code and the credential live next to each other and both are reachable from the same prompt, an injected instruction that convinces the agent to skip its own check, or to call the raw payment credential directly, defeats the limit entirely. A spending limit that the agent itself is responsible for enforcing is not really a limit on the agent. It's a suggestion to the agent.
Pink Agentic AI Payments enforces per-agent spending policies at the MCP layer, before a payment executes. Concretely:
The agent never holds a payment credential, only a Pink agent key. That key can ask Pink to check a policy, request a payment, or read the current budget. It cannot itself move money. When a payment request is allowed, Pink returns a single-use credential (in the sandbox, a test virtual card) that is locked to that specific payee and that specific amount, and expires 15 minutes after issuance. If the request should instead go to a human (for example, above an approval threshold), Pink returns pending_human and no credential is issued until a person approves it.
Policy decisions and spend recording happen as one step on Pink's side, not split between a check and a later update inside the agent. The "read balance, compare to limit, decide" sequence runs as part of evaluating the payment request itself, with no gap in which two concurrent requests can both read the same "under the limit" state. That's what we tested below.
Repeated requests with the same idempotency key return the original decision. An agent that retries a call (its own retry logic, a flaky network, a dropped response) doesn't get a second roll of the dice, and can't turn one approved payment into two by replaying it with a different amount under the same key.
We ran this against our public sandbox, which anyone can create a free workspace on with no real money involved. The setup: a startup template workspace, agent a_eng ("Eng Infra AI"), governed by a rule that reads "API credits: $200 a day per agent, then stop." We fired 40 parallel $7 payment requests at once, all to the same payee, all in the same instant.
curl -s -X POST https://agentic-sandbox.pinkwallet.com/v1/sandbox/workspaces \
-H 'content-type: application/json' \
-d '{"template":"startup","company":"race-test"}'
K=<a_eng's key from the response above>
for i in $(seq 1 40); do
curl -s -X POST https://agentic-sandbox.pinkwallet.com/v1/payments \
-H "Authorization: Bearer $K" -H 'content-type: application/json' \
-d '{"payee_id":"p_anth","amount":7,"purpose":"race test","local_hour":12}' &
done
wait
curl -s https://agentic-sandbox.pinkwallet.com/v1/budget -H "Authorization: Bearer $K"
40 requests at $7 each total $280, well over the $200/day rule. If the limit held, something close to 28 should be allowed ($196) and the rest blocked.
| Run | Requests | Allowed | Blocked | spent_today after |
|---|---|---|---|---|
| 1 | 40 x $7 | 28 | 12 | $196 |
| 2 | 40 x $7 | 28 | 12 | $196 |
Both runs, fired fresh against a new sandbox workspace each time, landed on the same split: 28 allowed, 12 blocked, and a final spent_today of $196, never above $200. No request saw a stale "you're still under the limit" answer after the limit was effectively reached. The full script, race_test.sh, is linked below; it creates its own workspace, so there's nothing to configure before running it.
A separate run with bigger payments makes the same point: 12 parallel $25 requests against the same $200/day rule came back 8 allowed, 4 blocked, spending exactly $200.
To be direct about the limits of this, because an agent payment system that overclaims is worse than one that's honest about gaps:
agentic-sandbox.pinkwallet.com with test credentials and no real money. Production access is by early access; nothing here claims production numbers.local_hour field is how you tell Pink what time it actually is where the payment is happening; some rules (like approval windows) are evaluated on the server's UTC clock unless you pass it, which can surprise you in testing if you skip it.
None of this is a guarantee that nothing can ever go wrong with an agent that can spend money; it's a description of where the current enforcement boundary sits and where it doesn't yet reach.
The sandbox is free, takes a POST to create a workspace, and the agent keys it hands back work immediately against both the REST API and an MCP endpoint with seven tools (pink.check_policy, pink.request_payment, pink.get_credential, pink.get_budget, pink.list_payees, pink.list_rules, pink.report_receipt).
./race_test.sh 40 7 to reproduce the table; it prints no secrets and creates its own disposable workspace)
If you've hit either of these failure modes building agent payments yourself, tell us how you handled it. And if you can make the script above overspend, we want to know.