{"slug": "the-app-started-guarding-before-the-invoice-arrived-an-ai-cost-cap-spend-guard", "title": "The App Started Guarding Before the Invoice Arrived — an AI Cost-Cap Spend Guard in AIO Helper", "summary": "A developer added a per-site monthly AI spend cap (default $50) to the API server of AIO Helper, a SEO analysis SaaS, so that requests whose estimated cost exceeds the remaining budget are blocked before the AI API is called. The mechanism, introduced on April 16, 2026, with an admin \"AI Budget\" tab added April 23, stores caps and usage records in the API server's database and covers the three text-generation features, though not embedding calls. The developer notes the strict cap guarantee under concurrency has not been verified against real Postgres and that after-the-fact recording cannot stop overruns from concurrent requests or costs above the estimate.", "body_md": "On April 16, 2026, I added a mechanism to the API server of AIO Helper, our SEO analysis SaaS, that stops processing before a site's monthly AI API spend exceeds its per-site cap (default $50). A request whose estimated cost exceeds the remaining budget is stopped before the AI is called. Usage is recorded after the AI call succeeds.\n\n`INSERT` instead of splitting them into two steps. That said, the strict cap guarantee under concurrency has not been verified against real Postgres. Also, actual costs above the estimate, or overruns caused by concurrent requests, cannot be stopped by after-the-fact recording.\nThe target is AIO Helper, our own SaaS for SEO operations. It ingests site data from sources such as Google Search Console and proposes improvements page by page. It has two parts: the admin UI that users operate, and the API server (`seo-api`) that does the processing. For an overview of the whole product, see [the AIO Helper introduction](https://dev.to/uehara/running-seo-improvements-as-one-flow-from-proposal-to-verified-effect-aio-helper-45o7).\n\nThe AI is called on the API server side. As of April 2026, three features called the text-generation AI API:\n\n`POST /v1/page-goals/auto-generate`). From a page's title and body, it drafts target keywords, the page's purpose, and the intended reader.`POST /v1/suggest`). Improvement proposals for one page.` POST /v1/suggest/batch`). Proposals for many pages, starting from those with the most clicks.\nThere were also calls that create search vectors (embeddings, which turn text into a list of numbers) to find related documents: the ones used inside the two suggestion features, and the API that rebuilds the vectors (`POST /v1/embeddings/rebuild`). The April mechanism did not cover these (explained later under \"What This Design Has Not Verified\").\n\nFor users, the value is that analysis finishes fast and manual research time goes down.\n\nBut an AI API costs money every time you call it. The bill grows not only when usage grows, but also when a bug repeats the same processing. Just showing the spend in an admin dashboard is not enough — by the time you notice, the cost has already been incurred.\n\nSo I built two things.\n\nOn April 16, a site's monthly cap was changed through the API (`PUT /v1/sites/:siteId/budget`); on April 23 I also added an \"AI Budget\" tab to the admin UI. The cap and the usage records live in the API server's database, and the three text-generation features read the same tables right before calling the AI.\n\nThe goal is not simply to stop at $50 a month. It is to be able to trace which feature spent how much, and to confirm whether costs actually went down after changing how we use it. I aimed for a state where \"stopping\" and \"improving\" both work from the same records.\n\nThe processing order is as follows.\n\nWhat matters in this order is not treating the stop decision as \"cleanup after the AI API returns an error.\" Once the request has gone out, you cannot cancel the cost. Only by stopping at the app's entry point does the cap function as a control.\n\nAt the same time, this record is not a substitute for the invoice. The app estimates from the token counts it knows and its own price table. It will not necessarily match the provider's final bill exactly. The in-app ledger is the number for stopping early; the invoice is the number you finally pay — separate roles.\n\nThe trigger was a test audit. The AIO Helper admin panel had a widget showing \"AI spend this month,\" but there were zero tests enforcing a cap. Even if you can calculate the spend, if you cannot stop before exceeding the cap, you keep calling pay-per-use generative AI models.\n\nThe first thing I did was separate the pricing calculation from the data model. I decided the price table would live statically in code, not be fetched from an external API. Three reasons.\n\n``` js\n// ai-cost.ts — unknown models are deliberately estimated high\nconst FALLBACK_PRICE = { input: 0.02, output: 0.08 };\n\n// calcCost() rounds to 6 digits, clamps negatives to 0, and never throws\n// estimateCostFromBytes() uses 3 bytes/token (between English ~4 and Japanese ~1.5-2)\n// and assumes output is 25% of input, for the preflight estimate\n```\n\nThe point is that `FALLBACK_PRICE` is deliberately high ($0.02/$0.08 per 1K). If you add a new model and forget to register it in the price table, its cost is treated as 0 and it slips past the cap. So treating unknown models as \"expensive\" is the safe side.\n\nThe core of the spend stop lives in `budget.ts`. There are two tables.\n\n```\n-- seo_site_budgets: per-site monthly cap (default $50)\n--   monthly_usd_cap NUMERIC(10,2)\n-- seo_ai_usage: append-only ledger, one row per call\n--   cost_usd NUMERIC(10,6), created_at timestamptz\n```\n\nI used `NUMERIC` instead of `FLOAT` for amounts to avoid rounding errors accumulating over millions of rows. What I agonized over most was how to handle two requests arriving right at the cap at the same time. Instead of `SELECT FOR UPDATE` with a transaction, I chose a single `INSERT` statement that includes the cap check.\n\n```\nINSERT INTO seo_ai_usage (site_id, model, endpoint, input_tokens, output_tokens, cost_usd)\nSELECT $1::text, $2::text, $3::text, $4::int, $5::int, $6::numeric\n WHERE (\n   COALESCE((SELECT monthly_usd_cap FROM seo_site_budgets WHERE site_id = $1), $7::numeric)\n   - COALESCE((\n       SELECT SUM(cost_usd) FROM seo_ai_usage\n        WHERE site_id = $1\n          AND created_at >= date_trunc('month', NOW())\n     ), 0)\n ) >= $6::numeric\nRETURNING id;\n```\n\n`$7` is the default cap ($50) used when the site has no budget row. The upfront remaining-budget check is done by `checkBudget()`, and if the budget is insufficient, it stops without calling the upstream AI fetch (the two single-page features return HTTP 402 Payment Required). `recordUsage()`, which runs after a successful AI call, returns `{ ok: false }` if this INSERT returns 0 rows. At that point the cost has already been incurred, so the request is not failed; it writes a warning log. Because nothing goes into the ledger, the full cost of that call is missing from the ledger, and since the recorded total does not grow, later cap checks do not reflect it either. `cap = 0` works as a shutoff setting meaning \"don't spend a cent.\"\n\nThe cap check is not a separate `SELECT` statement; it is a correlated subquery inside the `INSERT`'s `WHERE` clause. Within a single statement, no application-side work can slip in between the aggregation and the write.\n\nHowever, what I could verify independently was only the structure of this SQL. That two statements starting at the same time will always include each other's preceding rows in the aggregation has not been confirmed with contention tests against real Postgres. If you need a strict cap guarantee, you need to make per-site ordering explicit — locking the budget row, serializable isolation, or advisory locks.\n\nThis design also had a test environment problem. The tests ran on pg-mem (an in-memory Postgres), which did not implement `date_trunc('month', NOW())`, so according to the records at the time, 12 of the 18 tests failed with 500s.\n\n```\n12 out of 18 tests in integration-ai-cost-cap.test.ts failed\nOnly the 3 pure calcCost unit tests pass\n```\n\nAt the time of the record, the only tests passing were the `calcCost` unit tests that never touch the database. Some tests also failed for reasons other than the 500s; what the records show is a case where `PUT budget` returned 400 due to an input validation issue.\n\nAccording to the test records, registering `date_trunc` with `db.public.registerFunction()` cut the failures to 5. The rest were values inside the `INSERT...SELECT` whose types could not be resolved. node-postgres sends untyped values as strings, so I fixed it by adding explicit casts like `$1::text`.\n\nThe test records also include an experiment where I deliberately broke the `date_trunc` `WHERE` clause. 17 of the 18 tests still passed with the monthly extraction condition broken; only the \"don't count last month\" test failed. That test inserts a previous-month row dated 60 days back and confirms it is not aggregated. If someone accidentally deleted the monthly filter, this one test would be the only thing catching it. I learned that verification of a critical condition was concentrated in a single test.\n\nStopping at the cap alone means users suddenly get a 402 without knowing why it stopped. So I added a soft warning that sets a flag once the remaining budget drops below 10%.\n\n``` js\nexport const WARNING_REMAINING_PCT = 0.1;\nfunction isNearLimit(cap: number, remaining: number): boolean {\n  if (cap <= 0) return false;              // a kill switch is not a \"warning\"\n  return remaining / cap < WARNING_REMAINING_PCT;\n}\n```\n\nExplicitly excluding `cap <= 0` is so that \"usage is shut off\" and \"the cap is near\" are treated as separate notifications. According to the execution records, I first wrote 4 failing tests, and after implementing the warning condition, 240 tests passed. The change was committed as `feat: AI budget soft-warning flag at <10% remaining`.\n\nSeparating the warning from the shutoff also lets you separate the wording in the UI and notifications. If the remaining budget is low, the user can reduce their workload or ask an admin to raise the cap. If the cap is set to 0, an admin has intentionally stopped usage. Show both as the same yellow warning, and users will misread it as \"wait and it comes back.\"\n\nNotifying only that the cap was reached does not tell you what to fix next. For operations, you need at least these five things.\n\nThe April `seo_ai_usage` table has the columns (site, model, endpoint, token counts, cost, timestamp) to answer the first three. There was no screen yet for daily growth or before/after comparisons. Still, as long as each call is recorded, you can aggregate later and compare reductions with numbers instead of gut feeling.\n\nI cannot present this article as \"a finished version that never exceeds the budget even under concurrency.\" Four reasons.\n\nSo this implementation is not an endpoint. It is the first stage of treating cost as something to control. Follow-up articles move on to reconciling against billing data, identifying fixed costs, and re-measuring after changes. Skip this order, and you cannot explain reduction effects with numbers.\n\nThis cap control and ledger became the prototype for the measurement, caps, and records used in the cost improvements from May onward. The next month I investigated a problem where Cost Explorer returned $0, and later moved on to keeping \"measure → fix → re-measure\" on record. That way of thinking was already in place this April.\n\nOn the other hand, what this mechanism covers is AIO Helper's AI API spend. It cannot stop cloud-wide charges such as always-on Elastic Container Service or NAT Gateway fixed costs. Costs incurred outside the application need separate coverage with Cost Explorer and budget monitoring.\n\n`cap = 0` (stopped) into the \"almost there\" warning.", "url": "https://wpnews.pro/news/the-app-started-guarding-before-the-invoice-arrived-an-ai-cost-cap-spend-guard", "canonical_source": "https://dev.to/uehara/the-app-started-guarding-before-the-invoice-arrived-an-ai-cost-cap-spend-guard-in-aio-helper-4kgj", "published_at": "2026-10-06 01:01:00+00:00", "updated_at": "2026-10-06 01:17:29.306627+00:00", "lang": "en", "topics": ["ai-products", "ai-tools", "mlops", "ai-infrastructure"], "entities": ["AIO Helper", "Google Search Console", "Postgres"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/the-app-started-guarding-before-the-invoice-arrived-an-ai-cost-cap-spend-guard", "markdown": "https://wpnews.pro/news/the-app-started-guarding-before-the-invoice-arrived-an-ai-cost-cap-spend-guard.md", "text": "https://wpnews.pro/news/the-app-started-guarding-before-the-invoice-arrived-an-ai-cost-cap-spend-guard.txt", "jsonld": "https://wpnews.pro/news/the-app-started-guarding-before-the-invoice-arrived-an-ai-cost-cap-spend-guard.jsonld"}}