On April 16, 2026, I added a mechanism to the API server of AIO Helper, our SEO analysis SaaS, that stops processing before a site's monthly AI API spend exceeds its per-site cap (default $50). A request whose estimated cost exceeds the remaining budget is stopped before the AI is called. Usage is recorded after the AI call succeeds.
INSERT instead of splitting them into two steps. That said, the strict cap guarantee under concurrency has not been verified against real Postgres. Also, actual costs above the estimate, or overruns caused by concurrent requests, cannot be stopped by after-the-fact recording.
The target is AIO Helper, our own SaaS for SEO operations. It ingests site data from sources such as Google Search Console and proposes improvements page by page. It has two parts: the admin UI that users operate, and the API server (seo-api) that does the processing. For an overview of the whole product, see the AIO Helper introduction.
The AI is called on the API server side. As of April 2026, three features called the text-generation AI API:
POST /v1/page-goals/auto-generate). From a page's title and body, it drafts target keywords, the page's purpose, and the intended reader.POST /v1/suggest). Improvement proposals for one page. POST /v1/suggest/batch). Proposals for many pages, starting from those with the most clicks.
There were also calls that create search vectors (embeddings, which turn text into a list of numbers) to find related documents: the ones used inside the two suggestion features, and the API that rebuilds the vectors (POST /v1/embeddings/rebuild). The April mechanism did not cover these (explained later under "What This Design Has Not Verified").
For users, the value is that analysis finishes fast and manual research time goes down.
But an AI API costs money every time you call it. The bill grows not only when usage grows, but also when a bug repeats the same processing. Just showing the spend in an admin dashboard is not enough — by the time you notice, the cost has already been incurred.
So I built two things.
On April 16, a site's monthly cap was changed through the API (PUT /v1/sites/:siteId/budget); on April 23 I also added an "AI Budget" tab to the admin UI. The cap and the usage records live in the API server's database, and the three text-generation features read the same tables right before calling the AI.
The goal is not simply to stop at $50 a month. It is to be able to trace which feature spent how much, and to confirm whether costs actually went down after changing how we use it. I aimed for a state where "stopping" and "improving" both work from the same records.
The processing order is as follows.
What matters in this order is not treating the stop decision as "cleanup after the AI API returns an error." Once the request has gone out, you cannot cancel the cost. Only by stopping at the app's entry point does the cap function as a control.
At the same time, this record is not a substitute for the invoice. The app estimates from the token counts it knows and its own price table. It will not necessarily match the provider's final bill exactly. The in-app ledger is the number for stopping early; the invoice is the number you finally pay — separate roles.
The trigger was a test audit. The AIO Helper admin panel had a widget showing "AI spend this month," but there were zero tests enforcing a cap. Even if you can calculate the spend, if you cannot stop before exceeding the cap, you keep calling pay-per-use generative AI models.
The first thing I did was separate the pricing calculation from the data model. I decided the price table would live statically in code, not be fetched from an external API. Three reasons.
// ai-cost.ts — unknown models are deliberately estimated high
const FALLBACK_PRICE = { input: 0.02, output: 0.08 };
// calcCost() rounds to 6 digits, clamps negatives to 0, and never throws
// estimateCostFromBytes() uses 3 bytes/token (between English ~4 and Japanese ~1.5-2)
// and assumes output is 25% of input, for the preflight estimate
The point is that FALLBACK_PRICE is deliberately high ($0.02/$0.08 per 1K). If you add a new model and forget to register it in the price table, its cost is treated as 0 and it slips past the cap. So treating unknown models as "expensive" is the safe side.
The core of the spend stop lives in budget.ts. There are two tables.
-- seo_site_budgets: per-site monthly cap (default $50)
-- monthly_usd_cap NUMERIC(10,2)
-- seo_ai_usage: append-only ledger, one row per call
-- cost_usd NUMERIC(10,6), created_at timestamptz
I used NUMERIC instead of FLOAT for amounts to avoid rounding errors accumulating over millions of rows. What I agonized over most was how to handle two requests arriving right at the cap at the same time. Instead of SELECT FOR UPDATE with a transaction, I chose a single INSERT statement that includes the cap check.
INSERT INTO seo_ai_usage (site_id, model, endpoint, input_tokens, output_tokens, cost_usd)
SELECT $1::text, $2::text, $3::text, $4::int, $5::int, $6::numeric
WHERE (
COALESCE((SELECT monthly_usd_cap FROM seo_site_budgets WHERE site_id = $1), $7::numeric)
- COALESCE((
SELECT SUM(cost_usd) FROM seo_ai_usage
WHERE site_id = $1
AND created_at >= date_trunc('month', NOW())
), 0)
) >= $6::numeric
RETURNING id;
$7 is the default cap ($50) used when the site has no budget row. The upfront remaining-budget check is done by checkBudget(), and if the budget is insufficient, it stops without calling the upstream AI fetch (the two single-page features return HTTP 402 Payment Required). recordUsage(), which runs after a successful AI call, returns { ok: false } if this INSERT returns 0 rows. At that point the cost has already been incurred, so the request is not failed; it writes a warning log. Because nothing goes into the ledger, the full cost of that call is missing from the ledger, and since the recorded total does not grow, later cap checks do not reflect it either. cap = 0 works as a shutoff setting meaning "don't spend a cent."
The cap check is not a separate SELECT statement; it is a correlated subquery inside the INSERT's WHERE clause. Within a single statement, no application-side work can slip in between the aggregation and the write.
However, what I could verify independently was only the structure of this SQL. That two statements starting at the same time will always include each other's preceding rows in the aggregation has not been confirmed with contention tests against real Postgres. If you need a strict cap guarantee, you need to make per-site ordering explicit — locking the budget row, serializable isolation, or advisory locks.
This design also had a test environment problem. The tests ran on pg-mem (an in-memory Postgres), which did not implement date_trunc('month', NOW()), so according to the records at the time, 12 of the 18 tests failed with 500s.
12 out of 18 tests in integration-ai-cost-cap.test.ts failed
Only the 3 pure calcCost unit tests pass
At the time of the record, the only tests passing were the calcCost unit tests that never touch the database. Some tests also failed for reasons other than the 500s; what the records show is a case where PUT budget returned 400 due to an input validation issue.
According to the test records, registering date_trunc with db.public.registerFunction() cut the failures to 5. The rest were values inside the INSERT...SELECT whose types could not be resolved. node-postgres sends untyped values as strings, so I fixed it by adding explicit casts like $1::text.
The test records also include an experiment where I deliberately broke the date_trunc WHERE clause. 17 of the 18 tests still passed with the monthly extraction condition broken; only the "don't count last month" test failed. That test inserts a previous-month row dated 60 days back and confirms it is not aggregated. If someone accidentally deleted the monthly filter, this one test would be the only thing catching it. I learned that verification of a critical condition was concentrated in a single test.
Stopping at the cap alone means users suddenly get a 402 without knowing why it stopped. So I added a soft warning that sets a flag once the remaining budget drops below 10%.
export const WARNING_REMAINING_PCT = 0.1;
function isNearLimit(cap: number, remaining: number): boolean {
if (cap <= 0) return false; // a kill switch is not a "warning"
return remaining / cap < WARNING_REMAINING_PCT;
}
Explicitly excluding cap <= 0 is so that "usage is shut off" and "the cap is near" are treated as separate notifications. According to the execution records, I first wrote 4 failing tests, and after implementing the warning condition, 240 tests passed. The change was committed as feat: AI budget soft-warning flag at <10% remaining.
Separating the warning from the shutoff also lets you separate the wording in the UI and notifications. If the remaining budget is low, the user can reduce their workload or ask an admin to raise the cap. If the cap is set to 0, an admin has intentionally stopped usage. Show both as the same yellow warning, and users will misread it as "wait and it comes back."
Notifying only that the cap was reached does not tell you what to fix next. For operations, you need at least these five things.
The April seo_ai_usage table has the columns (site, model, endpoint, token counts, cost, timestamp) to answer the first three. There was no screen yet for daily growth or before/after comparisons. Still, as long as each call is recorded, you can aggregate later and compare reductions with numbers instead of gut feeling.
I cannot present this article as "a finished version that never exceeds the budget even under concurrency." Four reasons.
So this implementation is not an endpoint. It is the first stage of treating cost as something to control. Follow-up articles move on to reconciling against billing data, identifying fixed costs, and re-measuring after changes. Skip this order, and you cannot explain reduction effects with numbers.
This cap control and ledger became the prototype for the measurement, caps, and records used in the cost improvements from May onward. The next month I investigated a problem where Cost Explorer returned $0, and later moved on to keeping "measure → fix → re-measure" on record. That way of thinking was already in place this April.
On the other hand, what this mechanism covers is AIO Helper's AI API spend. It cannot stop cloud-wide charges such as always-on Elastic Container Service or NAT Gateway fixed costs. Costs incurred outside the application need separate coverage with Cost Explorer and budget monitoring.
cap = 0 (stopped) into the "almost there" warning.