A common first version of an LLM spending limit looks like this:
let used = 0;
async function call(prompt: string) {
if (used + ESTIMATE > LIMIT) throw new Error("quota exceeded"); // check
const res = await provider.complete(prompt); // call
used += res.usage.totalTokens; // add
return res;
}
It reads correctly. It has two defects, and both show up only under load or failure, which is when a spending limit matters.
Everything below is reproducible without an API key. The "provider" in the demo is a 5 ms timer. The code is in llm-quota-guard (MIT).
await
JavaScript is single-threaded, but await hands control back to the event loop. Ten requests arriving together all run the if before any of them reaches used += …:
limit=100, estimate=30, concurrent calls=10
naive counter : used=300 (limit exceeded by 200)
Each call estimated 30 units against a limit of 100. All ten passed the check, because used was still 0 when they asked. The limit was exceeded by a factor of three. Nothing was misconfigured; the order of operations is the problem.
The fix is to make the check and a claim on the budget one indivisible step, before the await. This is a reservation: a hold that counts against the limit immediately.
const hold = guard.reserve("tenant-42", 30); // throws if used + held + 30 > limit
With the same ten concurrent calls:
QuotaGuard : used=90, rejected=7
Three calls fit (3 × 30 = 90 ≤ 100). Seven are rejected before they reach the provider.
The opposite fix is to add the estimate before the call. That closes defect 1 and opens another: calls that fail never give the estimate back.
2) 4 calls that all fail (nothing was produced)
naive counter : used=120
QuotaGuard : used=0
Four failed calls consumed 120 units of a 100-unit budget while producing nothing. Failures often arrive in bursts (an upstream incident, a bad deploy, a rate-limit storm), so a burst can exhaust the budget while no useful work was done.
A reservation has two exits:
commit(actual) replaces the hold with what the provider actually billed.release() drops the hold and records nothing.
guard.run() wires both to the call's outcome: commit on success, release if the function throws.
You don't know the true cost before the call. The reservation uses your estimate, and commit(actual) trues it up afterwards. Two details matter:
overage, so you can see how good your estimates are. reserve() and commit() would otherwise lock quota forever, so each hold has a TTL (default 120 s). If a Two more rules remove ambiguity: settling twice returns the first result instead of counting twice (so a retry is safe), and usage belongs to the window in which the reservation was made, so a call that finishes after a window boundary doesn't charge the next window.
The limits are worth stating plainly:
Use provider-side hard limits as the backstop, and treat an in-app guard as the layer that gives your users and tenants predictable limits.
git clone https://github.com/soda4001/llm-quota-guard
cd llm-quota-guard
npm install
npm run example # prints the numbers above
npm test # 22 tests; they specify each rule described here
Developed and tested on Node 24 (CI runs the same commands). The implementation is about 320 lines including comments, with no runtime dependencies.
If you have handled this differently (Redis scripts, token buckets, provider-side budgets), counterexamples and corrections are welcome in the comments or as issues on the repo.
This article was written by an AI (Claude) under the direction of the account owner. It was not written or edited by hand. The numbers above are the output of the commands in the Reproduce section.