An agent calls a tool.
The tool times out.
So the agent tries again.
That sounds responsible.
Sometimes it is.
Sometimes it is a second charge on a customer's card.
A timeout tells you what the client saw.
It does not tell you what the server did.
That distinction matters a lot more now that agents send email, open tickets and move money.
This week a paper put real numbers on it.
None of the fixes are new.
Old idea.
New caller.
So let's build a tiny exactly-once layer.
By the end, you'll run one command:
npx tsx retry.ts
And watch four retry strategies hit four kinds of network fault, with a ledger counting every real charge.
No API key.
No real model.
Just TypeScript.
One honesty note: this is my small model of the paper's idea, not the Limbo benchmark. The faults are scripted. The billing service is fake.
Code: github.com/bobbyhalljr/tool-call-idempotency
A tool call is an intent with arguments.
A fake billing service commits charges to a ledger.
A fault injector breaks the first call in one of three ways.
Four strategies try to finish the job.
At the end, we count charges, not opinions.
It's also a small version of the architecture behind Roster: the agent proposes, the harness owns what actually happens.
The customer "acme", the $49.00 amount and the timings are example inputs.
You will need Node.js 18 or newer.
mkdir tool-call-idempotency
cd tool-call-idempotency
npm init -y
npm install --save-dev typescript tsx @types/node
Save the following blocks, in order, as retry.ts.
// retry.ts: a tiny exactly-once layer for agent tool calls.
// Everything is mocked: an in-memory billing service with injected network faults.
// No API key, no network, no real model. Customers, amounts and timings are example inputs.
import { createHash } from "node:crypto";
// Step 1: model the tool call
type ChargeArgs = { customer: string; amountCents: number };
type Charge = { id: string; customer: string; amountCents: number; at: number };
// lost-ack: the charge commits, the response is lost (client sees a timeout)
// late-commit: the request is still in flight when the client gives up; it lands 90s later
// redelivery: the transport delivers one request twice; the client sees one success
type Fault = "none" | "lost-ack" | "late-commit" | "redelivery";
type CallResult =
| { status: "ok"; charge: Charge }
| { status: "timeout" }
| { status: "error"; message: string };
Three faults. Three different lies.
lost-ack committed, but the client heard nothing.
late-commit is still in the pipe when the client gives up.
redelivery looks like a clean success. It charged twice.
From the client's side, the first two look identical.
// Step 2: a mock billing service that honors idempotency keys
class Billing {
ledger: Charge[] = [];
now = 0; // simulated seconds
private keys = new Map<string, { fingerprint: string; charge: Charge }>();
private inFlight: { at: number; args: ChargeArgs; key?: string }[] = [];
private calls = 0;
private seq = 0;
constructor(private fault: Fault) {}
private commit(args: ChargeArgs, key?: string): Charge | string {
const fingerprint = JSON.stringify(args);
if (key) {
const seen = this.keys.get(key);
if (seen && seen.fingerprint !== fingerprint) {
return "idempotency key reused with different parameters";
}
if (seen) return seen.charge; // replay: same key, same args, same charge
}
const charge = { id: `ch_${++this.seq}`, ...args, at: this.now };
this.ledger.push(charge);
if (key) this.keys.set(key, { fingerprint, charge });
return charge;
}
advance(seconds: number) {
this.now += seconds;
const due = this.inFlight.filter((r) => r.at <= this.now);
this.inFlight = this.inFlight.filter((r) => r.at > this.now);
for (const r of due) this.commit(r.args, r.key);
}
createCharge(args: ChargeArgs, key?: string): CallResult {
const first = ++this.calls === 1;
if (first && this.fault === "lost-ack") {
this.commit(args, key);
this.now += 30;
return { status: "timeout" };
}
if (first && this.fault === "late-commit") {
this.inFlight.push({ at: this.now + 90, args, key });
this.now += 30;
return { status: "timeout" };
}
const out = this.commit(args, key);
if (typeof out === "string") return { status: "error", message: out };
if (first && this.fault === "redelivery") this.commit(args, key);
return { status: "ok", charge: out };
}
listCharges(customer: string): Charge[] {
return this.ledger.filter((c) => c.customer === customer);
}
}
commit follows the Stripe rule.
Same key, same arguments: replay the first charge.
Same key, different arguments: refuse.
No key: every call is a new charge.
advance moves a fake clock, so a late commit can land after the client has moved on.
// Step 3: three retry strategies that look reasonable
type Strategy = (svc: Billing, args: ChargeArgs, taskId: string) => string;
const MAX_ATTEMPTS = 3;
const naiveRetry: Strategy = (svc, args) => {
for (let i = 1; i <= MAX_ATTEMPTS; i++) {
const res = svc.createCharge(args);
if (res.status === "ok") return "completed";
svc.advance(1);
}
return "failed";
};
const verifyThenRetry: Strategy = (svc, args) => {
for (let i = 1; i <= MAX_ATTEMPTS; i++) {
const res = svc.createCharge(args);
if (res.status === "ok") return "completed";
svc.advance(1);
const found = svc
.listCharges(args.customer)
.some((c) => c.amountCents === args.amountCents);
if (found) return "completed";
}
return "failed";
};
const newKeyPerAttempt: Strategy = (svc, args, taskId) => {
for (let i = 1; i <= MAX_ATTEMPTS; i++) {
const res = svc.createCharge(args, `${taskId}-retry${i}`);
if (res.status === "ok") return "completed";
svc.advance(1);
}
return "failed";
};
naiveRetry is what most of us write first.
verifyThenRetry checks the ledger before trying again. That is what the paper saw frontier models do.
newKeyPerAttempt uses keys, but makes a fresh one for every retry.
The paper calls this out by name. Every remaining duplicate in its keys-everywhere runs came from an agent that sent the first attempt without a key, or changed the key on retry, for example by appending -retry1.
A new key per attempt is not idempotency.
It is a receipt printer.
// Step 4: pin one key per intent, in the harness
function canonical(args: ChargeArgs): string {
return JSON.stringify(Object.fromEntries(Object.entries(args).sort()));
}
function intentKey(taskId: string, tool: string, args: ChargeArgs): string {
const hash = createHash("sha256").update(canonical(args)).digest("hex");
return `${taskId}:${tool}:${hash.slice(0, 12)}`;
}
const pinnedKey: Strategy = (svc, args, taskId) => {
const key = intentKey(taskId, "create_charge", args);
for (let i = 1; i <= MAX_ATTEMPTS; i++) {
const res = svc.createCharge(args, key);
if (res.status === "ok") return "completed";
if (res.status === "error") return `failed: ${res.message}`;
svc.advance(1);
}
return "unknown: escalate";
};
The key comes from the task, the tool and the canonical arguments.
Not from the model.
Not from the attempt number.
Retry ten times, and it is still the same key.
If the harness cannot get an answer, it says unknown: escalate instead of guessing.
The model can suggest a retry. The harness decides whether it is the same request.
// Step 5: run every strategy against every fault and count real charges
const strategies: [string, Strategy][] = [
["naive retry", naiveRetry],
["verify then retry", verifyThenRetry],
["new key per attempt", newKeyPerAttempt],
["pinned intent key", pinnedKey],
];
const faults: Fault[] = ["none", "lost-ack", "late-commit", "redelivery"];
const args: ChargeArgs = { customer: "acme", amountCents: 4900 };
console.log("Task: charge acme $49.00 exactly once (MOCK billing, no API key)\n");
let duplicates = 0;
let quietDuplicates = 0;
for (const fault of faults) {
console.log(`Fault: ${fault}`);
for (const [name, run] of strategies) {
const svc = new Billing(fault);
const report = run(svc, args, "task-42");
svc.advance(3600); // let anything still in flight land
const charges = svc.listCharges("acme").length;
const verdict = charges === 1 ? "ok" : "DUPLICATE";
if (charges > 1) duplicates++;
if (charges > 1 && report === "completed") quietDuplicates++;
console.log(
` ${name.padEnd(20)} charges=${charges} ${verdict.padEnd(9)} agent said: ${report}`,
);
}
console.log("");
}
console.log(`Duplicates: ${duplicates} of ${faults.length * strategies.length} runs`);
console.log(`Duplicates the agent reported as completed: ${quietDuplicates}\n`);
// Same key, different amount: the service refuses instead of guessing.
const svc = new Billing("none");
const key = intentKey("task-42", "create_charge", args);
svc.createCharge(args, key);
const changed = svc.createCharge({ customer: "acme", amountCents: 9900 }, key);
console.log("Key reuse with a different amount:");
console.log(` ${changed.status === "error" ? changed.message : "accepted"}`);
Run it:
npx tsx retry.ts
You should see:
Task: charge acme $49.00 exactly once (MOCK billing, no API key)
Fault: none
naive retry charges=1 ok agent said: completed
verify then retry charges=1 ok agent said: completed
new key per attempt charges=1 ok agent said: completed
pinned intent key charges=1 ok agent said: completed
Fault: lost-ack
naive retry charges=2 DUPLICATE agent said: completed
verify then retry charges=1 ok agent said: completed
new key per attempt charges=2 DUPLICATE agent said: completed
pinned intent key charges=1 ok agent said: completed
Fault: late-commit
naive retry charges=2 DUPLICATE agent said: completed
verify then retry charges=2 DUPLICATE agent said: completed
new key per attempt charges=2 DUPLICATE agent said: completed
pinned intent key charges=1 ok agent said: completed
Fault: redelivery
naive retry charges=2 DUPLICATE agent said: completed
verify then retry charges=2 DUPLICATE agent said: completed
new key per attempt charges=1 ok agent said: completed
pinned intent key charges=1 ok agent said: completed
Duplicates: 7 of 16 runs
Duplicates the agent reported as completed: 7
Key reuse with a different amount:
idempotency key reused with different parameters
Seven duplicates out of sixteen runs.
Every one of them reported completed.
Verify-then-retry wins lost-ack, then loses late-commit. The ledger check runs before the in-flight charge lands.
The pinned key is the only row that stays at one charge everywhere.
This is a teaching layer. Here is what a real one needs.
A key only works if the other side stores it. In the paper's native contract, only two of eleven non-idempotent write paths accepted one. The author notes that ratio, if anything, flatters many real APIs.
Stripe says keys can be pruned after they're at least 24 hours old. A retry after that window is a new request.
My ledger check matches on customer and amount. A real customer can legitimately have two $49 charges. Match on the key or a business ID, not on look-alike fields.
When there is no key and no reliable read path, the honest answer is "I don't know." Escalate. Don't flip a coin with someone's credit card.
The paper measured transparent client-side retries cutting exactly-once success from 72% to 50%. A retry the model never sees is a retry nobody reasons about.
My harness post said the model proposes and the harness decides.
This is the same split, one layer down.
Model ──→ "charge acme $49"
↓
Harness ──→ intent key = task + tool + args
↓
Tool ──→ seen this key? replay : execute
↓
Ledger ──→ exactly one charge
The model provides the intent.
The harness provides the identity.
The tool contract provides the replay.
The ledger provides the evidence.
The human provides the answer when the state is unknown.
Retries are free. Duplicates are not.
I'm building Roster around this idea: AI employees with real responsibilities, tools, memory, schedules and computer access. They work inside a lane, and every action they take lands in a log you can check.
If the same follow-ups, handoffs, and waiting loops keep eating your week, give them to an AI employee.