We fired 40 payments at our own agent at the same instant. 28 got through, 12 didn't. A developer at PinkWallet tested Pink Agentic AI Payments by firing 40 simultaneous payment requests at an agent with a $200 daily cap, and found 28 succeeded while 12 were blocked. The company says the system avoids the standard check-then-act race condition by enforcing per-agent spending policies at the MCP layer and issuing single-use, payee- and amount-locked credentials that expire after 15 minutes, so the agent never holds a real payment credential. 28 payments got through. 12 didn't. The cap was $200 a day, and we fired all 40 at the exact same moment on purpose, because we don't trust our own code more than we'd trust yours. Disclosure: I work on Pink at PinkWallet. Every number below came from a live run against our public sandbox, reproducible with the script linked at the bottom. That race is the easy failure mode to catch. The harder one: even a cap that holds doesn't matter if the agent is still sitting on a real payment credential it can use on its own. Pink Agentic AI Payments closes both, by keeping the check and the credential off the agent's side of the fence entirely. Here's what each failure looks like and how it's closed. A spending cap usually gets enforced like this, somewhere in application code or inside the agent's own tool call: read the amount spent so far, check it against the limit, and if it's under, let the payment go through. That works for one call at a time. It falls apart the moment an agent or a swarm of agents, or retried tool calls fires several payment requests in parallel. Here's why. Say the limit is $200/day and $190 has already been spent. Ten parallel calls each ask for $7. Each one reads "spent so far: $190," each one computes $190 + $7 = $197, which is under $200, so each one passes its own check. All ten go through. The agent just spent $70 against a $10 remaining budget, and nothing in that check-then-spend logic ever saw it coming, because the check and the spend were not the same atomic step. This isn't hypothetical. It's the standard "check-then-act" race condition, and it gets worse, not better, as agents get faster and start running tool calls concurrently to save latency. Even if you fix the race, a second problem remains: where does the actual payment credential live? In most setups, the agent process holds a real API key, card number, or wallet credential, and the "spending limit" is a piece of logic that runs in the same process, checking the same memory, before the agent is allowed to use that credential. The issue is that anything enforced inside the agent's own context can potentially be bypassed by what's inside that context, including text the agent reads from a tool result, a webpage, or a document, if that text is crafted to look like an instruction. If the limit-check code and the credential live next to each other and both are reachable from the same prompt, an injected instruction that convinces the agent to skip its own check, or to call the raw payment credential directly, defeats the limit entirely. A spending limit that the agent itself is responsible for enforcing is not really a limit on the agent. It's a suggestion to the agent. Pink Agentic AI Payments enforces per-agent spending policies at the MCP layer, before a payment executes. Concretely: The agent never holds a payment credential, only a Pink agent key. That key can ask Pink to check a policy, request a payment, or read the current budget. It cannot itself move money. When a payment request is allowed, Pink returns a single-use credential in the sandbox, a test virtual card that is locked to that specific payee and that specific amount, and expires 15 minutes after issuance. If the request should instead go to a human for example, above an approval threshold , Pink returns pending human and no credential is issued until a person approves it. Policy decisions and spend recording happen as one step on Pink's side, not split between a check and a later update inside the agent. The "read balance, compare to limit, decide" sequence runs as part of evaluating the payment request itself, with no gap in which two concurrent requests can both read the same "under the limit" state. That's what we tested below. Repeated requests with the same idempotency key return the original decision. An agent that retries a call its own retry logic, a flaky network, a dropped response doesn't get a second roll of the dice, and can't turn one approved payment into two by replaying it with a different amount under the same key. We ran this against our public sandbox, which anyone can create a free workspace on with no real money involved. The setup: a startup template workspace, agent a eng "Eng Infra AI" , governed by a rule that reads "API credits: $200 a day per agent, then stop." We fired 40 parallel $7 payment requests at once, all to the same payee, all in the same instant. create a disposable sandbox workspace curl -s -X POST https://agentic-sandbox.pinkwallet.com/v1/sandbox/workspaces \ -H 'content-type: application/json' \ -d '{"template":"startup","company":"race-test"}' - returns agents , including a eng with its own agent key K=