{"slug": "a-timeout-is-not-a-failure-build-idempotent-tool-calls-for-ai-agents-in", "title": "A Timeout Is Not a Failure: Build Idempotent Tool Calls for AI Agents in TypeScript", "summary": "A developer published a TypeScript demo showing how AI agents that retry timed-out tool calls can cause duplicate side effects like double-charging a customer's card, and built a small exactly-once layer using idempotency keys to prevent it. The mock billing service injects three fault types — lost-ack, late-commit and redelivery — and compares four retry strategies against a ledger that counts real charges, with the code available on GitHub.", "body_md": "An agent calls a tool.\n\nThe tool times out.\n\nSo the agent tries again.\n\nThat sounds responsible.\n\nSometimes it is.\n\nSometimes it is a second charge on a customer's card.\n\nA timeout tells you what the client saw.\n\nIt does not tell you what the server did.\n\nThat distinction matters a lot more now that agents send email, open tickets and move money.\n\nThis week a paper put real numbers on it.\n\nNone of the fixes are new.\n\nOld idea.\n\nNew caller.\n\nSo let's build a tiny exactly-once layer.\n\nBy the end, you'll run one command:\n\n```\nnpx tsx retry.ts\n```\n\nAnd watch four retry strategies hit four kinds of network fault, with a ledger counting every real charge.\n\nNo API key.\n\nNo real model.\n\nJust TypeScript.\n\nOne honesty note: this is my small model of the paper's idea, not the Limbo benchmark. The faults are scripted. The billing service is fake.\n\n**Code:** [github.com/bobbyhalljr/tool-call-idempotency](https://github.com/bobbyhalljr/tool-call-idempotency)\n\nA tool call is an intent with arguments.\n\nA fake billing service commits charges to a ledger.\n\nA fault injector breaks the first call in one of three ways.\n\nFour strategies try to finish the job.\n\nAt the end, we count charges, not opinions.\n\nIt's also a small version of the architecture behind [Roster](https://get-roster.com): the agent proposes, the harness owns what actually happens.\n\nThe customer \"acme\", the $49.00 amount and the timings are example inputs.\n\nYou will need Node.js 18 or newer.\n\n```\nmkdir tool-call-idempotency\ncd tool-call-idempotency\n\nnpm init -y\nnpm install --save-dev typescript tsx @types/node\n```\n\nSave the following blocks, in order, as `retry.ts`.\n\n```\n// retry.ts: a tiny exactly-once layer for agent tool calls.\n// Everything is mocked: an in-memory billing service with injected network faults.\n// No API key, no network, no real model. Customers, amounts and timings are example inputs.\nimport { createHash } from \"node:crypto\";\n\n// Step 1: model the tool call\ntype ChargeArgs = { customer: string; amountCents: number };\ntype Charge = { id: string; customer: string; amountCents: number; at: number };\n\n// lost-ack:    the charge commits, the response is lost (client sees a timeout)\n// late-commit: the request is still in flight when the client gives up; it lands 90s later\n// redelivery:  the transport delivers one request twice; the client sees one success\ntype Fault = \"none\" | \"lost-ack\" | \"late-commit\" | \"redelivery\";\n\ntype CallResult =\n  | { status: \"ok\"; charge: Charge }\n  | { status: \"timeout\" }\n  | { status: \"error\"; message: string };\n```\n\nThree faults. Three different lies.\n\n`lost-ack` committed, but the client heard nothing.\n\n`late-commit` is still in the pipe when the client gives up.\n\n`redelivery` looks like a clean success. It charged twice.\n\n**From the client's side, the first two look identical.**\n\n```\n// Step 2: a mock billing service that honors idempotency keys\nclass Billing {\n  ledger: Charge[] = [];\n  now = 0; // simulated seconds\n  private keys = new Map<string, { fingerprint: string; charge: Charge }>();\n  private inFlight: { at: number; args: ChargeArgs; key?: string }[] = [];\n  private calls = 0;\n  private seq = 0;\n\n  constructor(private fault: Fault) {}\n\n  private commit(args: ChargeArgs, key?: string): Charge | string {\n    const fingerprint = JSON.stringify(args);\n    if (key) {\n      const seen = this.keys.get(key);\n      if (seen && seen.fingerprint !== fingerprint) {\n        return \"idempotency key reused with different parameters\";\n      }\n      if (seen) return seen.charge; // replay: same key, same args, same charge\n    }\n    const charge = { id: `ch_${++this.seq}`, ...args, at: this.now };\n    this.ledger.push(charge);\n    if (key) this.keys.set(key, { fingerprint, charge });\n    return charge;\n  }\n\n  advance(seconds: number) {\n    this.now += seconds;\n    const due = this.inFlight.filter((r) => r.at <= this.now);\n    this.inFlight = this.inFlight.filter((r) => r.at > this.now);\n    for (const r of due) this.commit(r.args, r.key);\n  }\n\n  createCharge(args: ChargeArgs, key?: string): CallResult {\n    const first = ++this.calls === 1;\n    if (first && this.fault === \"lost-ack\") {\n      this.commit(args, key);\n      this.now += 30;\n      return { status: \"timeout\" };\n    }\n    if (first && this.fault === \"late-commit\") {\n      this.inFlight.push({ at: this.now + 90, args, key });\n      this.now += 30;\n      return { status: \"timeout\" };\n    }\n    const out = this.commit(args, key);\n    if (typeof out === \"string\") return { status: \"error\", message: out };\n    if (first && this.fault === \"redelivery\") this.commit(args, key);\n    return { status: \"ok\", charge: out };\n  }\n\n  listCharges(customer: string): Charge[] {\n    return this.ledger.filter((c) => c.customer === customer);\n  }\n}\n```\n\n`commit` follows the Stripe rule.\n\nSame key, same arguments: replay the first charge.\n\nSame key, different arguments: refuse.\n\nNo key: every call is a new charge.\n\n`advance` moves a fake clock, so a late commit can land after the client has moved on.\n\n```\n// Step 3: three retry strategies that look reasonable\ntype Strategy = (svc: Billing, args: ChargeArgs, taskId: string) => string;\nconst MAX_ATTEMPTS = 3;\n\nconst naiveRetry: Strategy = (svc, args) => {\n  for (let i = 1; i <= MAX_ATTEMPTS; i++) {\n    const res = svc.createCharge(args);\n    if (res.status === \"ok\") return \"completed\";\n    svc.advance(1);\n  }\n  return \"failed\";\n};\n\nconst verifyThenRetry: Strategy = (svc, args) => {\n  for (let i = 1; i <= MAX_ATTEMPTS; i++) {\n    const res = svc.createCharge(args);\n    if (res.status === \"ok\") return \"completed\";\n    svc.advance(1);\n    const found = svc\n      .listCharges(args.customer)\n      .some((c) => c.amountCents === args.amountCents);\n    if (found) return \"completed\";\n  }\n  return \"failed\";\n};\n\nconst newKeyPerAttempt: Strategy = (svc, args, taskId) => {\n  for (let i = 1; i <= MAX_ATTEMPTS; i++) {\n    const res = svc.createCharge(args, `${taskId}-retry${i}`);\n    if (res.status === \"ok\") return \"completed\";\n    svc.advance(1);\n  }\n  return \"failed\";\n};\n```\n\n`naiveRetry` is what most of us write first.\n\n`verifyThenRetry` checks the ledger before trying again. That is what the paper saw frontier models do.\n\n`newKeyPerAttempt` uses keys, but makes a fresh one for every retry.\n\nThe paper calls this out by name. Every remaining duplicate in its keys-everywhere runs came from an agent that sent the first attempt without a key, or changed the key on retry, for example by appending `-retry1`.\n\nA new key per attempt is not idempotency.\n\nIt is a receipt printer.\n\n```\n// Step 4: pin one key per intent, in the harness\nfunction canonical(args: ChargeArgs): string {\n  return JSON.stringify(Object.fromEntries(Object.entries(args).sort()));\n}\n\nfunction intentKey(taskId: string, tool: string, args: ChargeArgs): string {\n  const hash = createHash(\"sha256\").update(canonical(args)).digest(\"hex\");\n  return `${taskId}:${tool}:${hash.slice(0, 12)}`;\n}\n\nconst pinnedKey: Strategy = (svc, args, taskId) => {\n  const key = intentKey(taskId, \"create_charge\", args);\n  for (let i = 1; i <= MAX_ATTEMPTS; i++) {\n    const res = svc.createCharge(args, key);\n    if (res.status === \"ok\") return \"completed\";\n    if (res.status === \"error\") return `failed: ${res.message}`;\n    svc.advance(1);\n  }\n  return \"unknown: escalate\";\n};\n```\n\nThe key comes from the task, the tool and the canonical arguments.\n\nNot from the model.\n\nNot from the attempt number.\n\nRetry ten times, and it is still the same key.\n\nIf the harness cannot get an answer, it says `unknown: escalate` instead of guessing.\n\n**The model can suggest a retry. The harness decides whether it is the same request.**\n\n``` js\n// Step 5: run every strategy against every fault and count real charges\nconst strategies: [string, Strategy][] = [\n  [\"naive retry\", naiveRetry],\n  [\"verify then retry\", verifyThenRetry],\n  [\"new key per attempt\", newKeyPerAttempt],\n  [\"pinned intent key\", pinnedKey],\n];\nconst faults: Fault[] = [\"none\", \"lost-ack\", \"late-commit\", \"redelivery\"];\nconst args: ChargeArgs = { customer: \"acme\", amountCents: 4900 };\n\nconsole.log(\"Task: charge acme $49.00 exactly once (MOCK billing, no API key)\\n\");\nlet duplicates = 0;\nlet quietDuplicates = 0;\nfor (const fault of faults) {\n  console.log(`Fault: ${fault}`);\n  for (const [name, run] of strategies) {\n    const svc = new Billing(fault);\n    const report = run(svc, args, \"task-42\");\n    svc.advance(3600); // let anything still in flight land\n    const charges = svc.listCharges(\"acme\").length;\n    const verdict = charges === 1 ? \"ok\" : \"DUPLICATE\";\n    if (charges > 1) duplicates++;\n    if (charges > 1 && report === \"completed\") quietDuplicates++;\n    console.log(\n      `  ${name.padEnd(20)} charges=${charges}  ${verdict.padEnd(9)}  agent said: ${report}`,\n    );\n  }\n  console.log(\"\");\n}\nconsole.log(`Duplicates: ${duplicates} of ${faults.length * strategies.length} runs`);\nconsole.log(`Duplicates the agent reported as completed: ${quietDuplicates}\\n`);\n\n// Same key, different amount: the service refuses instead of guessing.\nconst svc = new Billing(\"none\");\nconst key = intentKey(\"task-42\", \"create_charge\", args);\nsvc.createCharge(args, key);\nconst changed = svc.createCharge({ customer: \"acme\", amountCents: 9900 }, key);\nconsole.log(\"Key reuse with a different amount:\");\nconsole.log(`  ${changed.status === \"error\" ? changed.message : \"accepted\"}`);\n```\n\nRun it:\n\n```\nnpx tsx retry.ts\n```\n\nYou should see:\n\n```\nTask: charge acme $49.00 exactly once (MOCK billing, no API key)\n\nFault: none\n  naive retry          charges=1  ok         agent said: completed\n  verify then retry    charges=1  ok         agent said: completed\n  new key per attempt  charges=1  ok         agent said: completed\n  pinned intent key    charges=1  ok         agent said: completed\n\nFault: lost-ack\n  naive retry          charges=2  DUPLICATE  agent said: completed\n  verify then retry    charges=1  ok         agent said: completed\n  new key per attempt  charges=2  DUPLICATE  agent said: completed\n  pinned intent key    charges=1  ok         agent said: completed\n\nFault: late-commit\n  naive retry          charges=2  DUPLICATE  agent said: completed\n  verify then retry    charges=2  DUPLICATE  agent said: completed\n  new key per attempt  charges=2  DUPLICATE  agent said: completed\n  pinned intent key    charges=1  ok         agent said: completed\n\nFault: redelivery\n  naive retry          charges=2  DUPLICATE  agent said: completed\n  verify then retry    charges=2  DUPLICATE  agent said: completed\n  new key per attempt  charges=1  ok         agent said: completed\n  pinned intent key    charges=1  ok         agent said: completed\n\nDuplicates: 7 of 16 runs\nDuplicates the agent reported as completed: 7\n\nKey reuse with a different amount:\n  idempotency key reused with different parameters\n```\n\nSeven duplicates out of sixteen runs.\n\nEvery one of them reported `completed`.\n\nVerify-then-retry wins `lost-ack`, then loses `late-commit`. The ledger check runs before the in-flight charge lands.\n\nThe pinned key is the only row that stays at one charge everywhere.\n\nThis is a teaching layer. Here is what a real one needs.\n\nA key only works if the other side stores it. In the paper's native contract, only two of eleven non-idempotent write paths accepted one. The author notes that ratio, if anything, flatters many real APIs.\n\nStripe says keys can be pruned after they're at least 24 hours old. A retry after that window is a new request.\n\nMy ledger check matches on customer and amount. A real customer can legitimately have two $49 charges. Match on the key or a business ID, not on look-alike fields.\n\nWhen there is no key and no reliable read path, the honest answer is \"I don't know.\" Escalate. Don't flip a coin with someone's credit card.\n\nThe paper measured transparent client-side retries cutting exactly-once success from 72% to 50%. A retry the model never sees is a retry nobody reasons about.\n\nMy [harness post](https://dev.to/bobbyhalljr/what-is-an-agent-harness-build-a-tiny-one-in-typescript-4fde) said the model proposes and the harness decides.\n\nThis is the same split, one layer down.\n\n```\nModel ──→ \"charge acme $49\"\n             ↓\nHarness ──→ intent key = task + tool + args\n             ↓\nTool ──→ seen this key? replay : execute\n             ↓\nLedger ──→ exactly one charge\n```\n\nThe model provides the intent.\n\nThe harness provides the identity.\n\nThe tool contract provides the replay.\n\nThe ledger provides the evidence.\n\nThe human provides the answer when the state is unknown.\n\n**Retries are free. Duplicates are not.**\n\nI'm building Roster around this idea: AI employees with real responsibilities, tools, memory, schedules and computer access. They work inside a lane, and every action they take lands in a log you can check.\n\nIf the same follow-ups, handoffs, and waiting loops keep eating your week, give them to an AI employee.", "url": "https://wpnews.pro/news/a-timeout-is-not-a-failure-build-idempotent-tool-calls-for-ai-agents-in", "canonical_source": "https://dev.to/bobbyhalljr/a-timeout-is-not-a-failure-build-idempotent-tool-calls-for-ai-agents-in-typescript-54d1", "published_at": "2026-10-04 13:06:59+00:00", "updated_at": "2026-10-04 13:12:54.423229+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools"], "entities": ["Roster", "GitHub", "TypeScript", "Node.js"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/a-timeout-is-not-a-failure-build-idempotent-tool-calls-for-ai-agents-in", "markdown": "https://wpnews.pro/news/a-timeout-is-not-a-failure-build-idempotent-tool-calls-for-ai-agents-in.md", "text": "https://wpnews.pro/news/a-timeout-is-not-a-failure-build-idempotent-tool-calls-for-ai-agents-in.txt", "jsonld": "https://wpnews.pro/news/a-timeout-is-not-a-failure-build-idempotent-tool-calls-for-ai-agents-in.jsonld"}}