{"slug": "two-ways-a-simple-llm-token-counter-goes-wrong-with-a-demo-you-can-run", "title": "Two ways a simple LLM token counter goes wrong (with a demo you can run)", "summary": "A developer published llm-quota-guard, an MIT-licensed Node library that fixes two concurrency defects in naive LLM token-budget counters: concurrent requests all passing the same pre-await check, and failed calls permanently consuming estimated quota. The library replaces the check-then-call-then-add pattern with a reserve/commit/release reservation model, showing that ten concurrent 30-unit calls against a 100-unit limit drop from 300 units used to 90 used with 7 rejections, and four failed calls consume 0 instead of 120. The roughly 320-line implementation has no runtime dependencies and ships with 22 tests, with the author inviting counterexamples from users of Redis scripts, token buckets, or provider-side budgets.", "body_md": "A common first version of an LLM spending limit looks like this:\n\n``` js\nlet used = 0;\n\nasync function call(prompt: string) {\n  if (used + ESTIMATE > LIMIT) throw new Error(\"quota exceeded\"); // check\n  const res = await provider.complete(prompt);                    // call\n  used += res.usage.totalTokens;                                  // add\n  return res;\n}\n```\n\nIt reads correctly. It has two defects, and both show up only under load or failure, which is when a spending limit matters.\n\nEverything below is reproducible without an API key. The \"provider\" in the demo is a 5 ms timer. The code is in [`llm-quota-guard`](https://github.com/soda4001/llm-quota-guard) (MIT).\n\n`await`\nJavaScript is single-threaded, but `await` hands control back to the event loop. Ten requests arriving together all run the `if` before any of them reaches `used += …`:\n\n```\nlimit=100, estimate=30, concurrent calls=10\nnaive counter : used=300 (limit exceeded by 200)\n```\n\nEach call estimated 30 units against a limit of 100. All ten passed the check, because `used` was still 0 when they asked. The limit was exceeded by a factor of three. Nothing was misconfigured; the order of operations is the problem.\n\nThe fix is to make the check and a *claim on the budget* one indivisible step, before the `await`. This is a **reservation**: a hold that counts against the limit immediately.\n\n``` js\nconst hold = guard.reserve(\"tenant-42\", 30); // throws if used + held + 30 > limit\n```\n\nWith the same ten concurrent calls:\n\n```\nQuotaGuard    : used=90, rejected=7\n```\n\nThree calls fit (3 × 30 = 90 ≤ 100). Seven are rejected before they reach the provider.\n\nThe opposite fix is to add the estimate *before* the call. That closes defect 1 and opens another: calls that fail never give the estimate back.\n\n```\n2) 4 calls that all fail (nothing was produced)\n   naive counter : used=120\n   QuotaGuard    : used=0\n```\n\nFour failed calls consumed 120 units of a 100-unit budget while producing nothing. Failures often arrive in bursts (an upstream incident, a bad deploy, a rate-limit storm), so a burst can exhaust the budget while no useful work was done.\n\nA reservation has two exits:\n\n`commit(actual)` replaces the hold with what the provider actually billed.`release()` drops the hold and records nothing.\n`guard.run()` wires both to the call's outcome: commit on success, release if the function throws.\n\nYou don't know the true cost before the call. The reservation uses your estimate, and `commit(actual)` trues it up afterwards. Two details matter:\n\n`overage`, so you can see how good your estimates are.` reserve()` and `commit()` would otherwise lock quota forever, so each hold has a TTL (default 120 s). If a Two more rules remove ambiguity: settling twice returns the first result instead of counting twice (so a retry is safe), and usage belongs to the window in which the reservation was made, so a call that finishes after a window boundary doesn't charge the next window.\n\nThe limits are worth stating plainly:\n\nUse provider-side hard limits as the backstop, and treat an in-app guard as the layer that gives *your* users and tenants predictable limits.\n\n```\ngit clone https://github.com/soda4001/llm-quota-guard\ncd llm-quota-guard\nnpm install\nnpm run example   # prints the numbers above\nnpm test          # 22 tests; they specify each rule described here\n```\n\nDeveloped and tested on Node 24 (CI runs the same commands). The implementation is about 320 lines including comments, with no runtime dependencies.\n\nIf you have handled this differently (Redis scripts, token buckets, provider-side budgets), counterexamples and corrections are welcome in the comments or as issues on the repo.\n\n*This article was written by an AI (Claude) under the direction of the account owner. It was not written or edited by hand. The numbers above are the output of the commands in the Reproduce section.*", "url": "https://wpnews.pro/news/two-ways-a-simple-llm-token-counter-goes-wrong-with-a-demo-you-can-run", "canonical_source": "https://dev.to/soda4001/two-ways-a-simple-llm-token-counter-goes-wrong-with-a-demo-you-can-run-1134", "published_at": "2026-10-04 01:59:32+00:00", "updated_at": "2026-10-04 02:07:43.269812+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "developer-tools", "mlops"], "entities": ["llm-quota-guard", "Node.js", "Claude", "GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/two-ways-a-simple-llm-token-counter-goes-wrong-with-a-demo-you-can-run", "markdown": "https://wpnews.pro/news/two-ways-a-simple-llm-token-counter-goes-wrong-with-a-demo-you-can-run.md", "text": "https://wpnews.pro/news/two-ways-a-simple-llm-token-counter-goes-wrong-with-a-demo-you-can-run.txt", "jsonld": "https://wpnews.pro/news/two-ways-a-simple-llm-token-counter-goes-wrong-with-a-demo-you-can-run.jsonld"}}