cd /news/large-language-models/two-ways-a-simple-llm-token-counter-… · home › topics › large-language-models › article
[ARTICLE · art-144666] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Two ways a simple LLM token counter goes wrong (with a demo you can run)

A developer published llm-quota-guard, an MIT-licensed Node library that fixes two concurrency defects in naive LLM token-budget counters: concurrent requests all passing the same pre-await check, and failed calls permanently consuming estimated quota. The library replaces the check-then-call-then-add pattern with a reserve/commit/release reservation model, showing that ten concurrent 30-unit calls against a 100-unit limit drop from 300 units used to 90 used with 7 rejections, and four failed calls consume 0 instead of 120. The roughly 320-line implementation has no runtime dependencies and ships with 22 tests, with the author inviting counterexamples from users of Redis scripts, token buckets, or provider-side budgets.

by read3 min views2 publishedOct 4, 2026

A common first version of an LLM spending limit looks like this:

let used = 0;

async function call(prompt: string) {
  if (used + ESTIMATE > LIMIT) throw new Error("quota exceeded"); // check
  const res = await provider.complete(prompt);                    // call
  used += res.usage.totalTokens;                                  // add
  return res;
}

It reads correctly. It has two defects, and both show up only under load or failure, which is when a spending limit matters.

Everything below is reproducible without an API key. The "provider" in the demo is a 5 ms timer. The code is in llm-quota-guard (MIT).

await JavaScript is single-threaded, but await hands control back to the event loop. Ten requests arriving together all run the if before any of them reaches used += …:

limit=100, estimate=30, concurrent calls=10
naive counter : used=300 (limit exceeded by 200)

Each call estimated 30 units against a limit of 100. All ten passed the check, because used was still 0 when they asked. The limit was exceeded by a factor of three. Nothing was misconfigured; the order of operations is the problem.

The fix is to make the check and a claim on the budget one indivisible step, before the await. This is a reservation: a hold that counts against the limit immediately.

const hold = guard.reserve("tenant-42", 30); // throws if used + held + 30 > limit

With the same ten concurrent calls:

QuotaGuard    : used=90, rejected=7

Three calls fit (3 × 30 = 90 ≤ 100). Seven are rejected before they reach the provider.

The opposite fix is to add the estimate before the call. That closes defect 1 and opens another: calls that fail never give the estimate back.

2) 4 calls that all fail (nothing was produced)
   naive counter : used=120
   QuotaGuard    : used=0

Four failed calls consumed 120 units of a 100-unit budget while producing nothing. Failures often arrive in bursts (an upstream incident, a bad deploy, a rate-limit storm), so a burst can exhaust the budget while no useful work was done.

A reservation has two exits:

commit(actual) replaces the hold with what the provider actually billed.release() drops the hold and records nothing. guard.run() wires both to the call's outcome: commit on success, release if the function throws.

You don't know the true cost before the call. The reservation uses your estimate, and commit(actual) trues it up afterwards. Two details matter:

overage, so you can see how good your estimates are. reserve() and commit() would otherwise lock quota forever, so each hold has a TTL (default 120 s). If a Two more rules remove ambiguity: settling twice returns the first result instead of counting twice (so a retry is safe), and usage belongs to the window in which the reservation was made, so a call that finishes after a window boundary doesn't charge the next window.

The limits are worth stating plainly:

Use provider-side hard limits as the backstop, and treat an in-app guard as the layer that gives your users and tenants predictable limits.

git clone https://github.com/soda4001/llm-quota-guard
cd llm-quota-guard
npm install
npm run example   # prints the numbers above
npm test          # 22 tests; they specify each rule described here

Developed and tested on Node 24 (CI runs the same commands). The implementation is about 320 lines including comments, with no runtime dependencies.

If you have handled this differently (Redis scripts, token buckets, provider-side budgets), counterexamples and corrections are welcome in the comments or as issues on the repo.

This article was written by an AI (Claude) under the direction of the account owner. It was not written or edited by hand. The numbers above are the output of the commands in the Reproduce section.

── more in #large-language-models 4 stories · sorted by recency
── more on @llm-quota-guard 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/two-ways-a-simple-ll…] indexed:0 read:3min 2026-10-04 · —