{"slug": "model-calls-are-metered-most-ai-features-are-built-like-they-are-free", "title": "Model Calls Are Metered. Most AI Features Are Built Like They Are Free.", "summary": "A developer who builds client voice agents and subscription products argues that AI features should treat model calls as a metered resource from the first sprint, since hosted models charge per request by tokens in and out and usage is uneven, input size is invisible, and agent loops multiply calls. The recommended practices include logging every model call with feature, tenant, model, token counts and whether the result was used, tracking cost per business unit such as call minute or active user, generating summaries and classifications on change rather than on view, exact-match and prompt caching, and batching non-urgent work.", "body_md": "Most of what we build at work has a cost profile that barely moves. A server, a database and a handful of SaaS seats cost about the same on a quiet Tuesday as on launch day. Then you ship your first AI feature, and for the first time a single click has a price attached.\n\nThat price rarely shows up in the demo. It shows up on the first full invoice after launch, when a feature that cost almost nothing in testing is being called far more often than anyone assumed, with far more context than anyone intended. The usual reaction is to shop for a cheaper model. Sometimes that helps. More often the bill is high because of how the feature is built, and a cheaper model just lowers the unit price of a wasteful design.\n\nSo this is about the design. These are the habits I now build in from the first sprint, on client voice agents and on my own products.\n\nHosted models charge per request, priced by tokens in and tokens out. Three things follow from that, and all three bite teams who are used to fixed infrastructure.\n\n**Usage is uneven.** Your heaviest users do not cost a bit more than your lightest. They can cost many times more, because they trigger the feature more often and usually on bigger inputs. An average cost per user hides exactly the accounts that hurt.\n\n**Input size is invisible.** Nobody sees the prompt. A feature that sends a whole document, the full chat history and a long instruction block on every request looks identical in the UI to one that sends a paragraph. The difference only appears on the bill.\n\n**Loops multiply.** An agent that calls a model, reads the result and calls again can make many calls for one user action. A retry bug, or an agent that never decides it is done, can make thousands.\n\nNone of this is a reason to avoid AI features. It is a reason to treat model calls as a metered resource from day one, the same way you already treat SMS or payment fees.\n\nA monthly total tells you almost nothing. The number that matters is cost per unit of whatever the business earns on: per call handled, per document processed, per active user.\n\nThat means every model call gets logged with enough attached to answer \"what was this for\". Something roughly like this:\n\n```\n// illustrative shape, one row per model call\ntype ModelCallLog = {\n  feature: \"summary\" | \"classify\" | \"agent_step\";\n  tenantId: string;\n  model: string;\n  inputTokens: number;\n  outputTokens: number;\n  used: boolean; // did the user keep the result, or regenerate it?\n  at: Date;\n};\n```\n\nIt is a small amount of engineering, and without it every later decision is a guess. With it you can answer the questions that actually drive spend: which feature is most of the bill, what the top tenth of users cost against the median, and how much you pay for results people throw away.\n\nPick the unit that matches revenue. On voice agents, like the ones we built for CallGuard AI and CallSetter AI, the natural unit is the call minute, because that is how those businesses think about their own economics. On a subscription product it is the active user per month, because that is what the subscription pays for.\n\nThis is the single biggest saving in most AI products.\n\n**Generate on change, not on view.** If a record gets a summary, generate it when the record changes and store it. Do not regenerate it every time someone opens the page. Read many times, written once, should cost one call. Same for tags, classifications, extracted fields and embeddings.\n\n**Cache what repeats.** Identical requests are more common than you think, and exact-match caching is simple and safe. Most major providers also offer prompt caching for a repeated long prefix, so put the stable part of the prompt (instructions, a manual, a policy) first and the variable part last. Same behavior, lower cost.\n\n**Batch what is not urgent.** Overnight reprocessing, bulk classification and backfills can often go through a provider's batch interface, which is usually priced below real time. That is a scheduling decision, not a model decision.\n\nI learned the caching one on my own product. In Upwork Scout, the AI fit score is stored per user and job, so a job is never scored twice for the same person. That one decision detached model spend from how often the scanner runs. I wrote up the whole matcher in [Cheap Filters First, LLM Last](https://dev.to/nabeelbaghoor/cheap-filters-first-llm-last-running-an-ai-matcher-inside-a-cron-job-702), so I will not repeat it here.\n\nContext is the quiet cost driver. Teams pass in everything because deciding what the task needs is more work, and the model copes. It just copes expensively.\n\nSmaller inputs are also faster, which matters a lot in anything real time, like a voice agent where a caller is sitting in silence waiting.\n\nRequests are not equally hard. A large share are routine: classify this message, extract these fields, answer from the provided text. A small model often handles those fine. A smaller share genuinely need a frontier model.\n\nThe pattern I use: for each task, find the cheapest model that passes your own eval set, send the task there, and give it an escalation path when it signals it cannot do the job. That only works if switching models does not mean rewriting the feature, so keep a thin layer between product code and the provider SDK. Even something as small as reading the model name from an env var instead of hardcoding it buys you that.\n\nAnd sometimes the right model is no model. If a rule, a lookup or a regex does the job reliably, it is cheaper, faster and far easier to test.\n\nMonitoring tells you about a problem after it happened. Limits stop it from getting expensive while it happens.\n\n**Per user and per plan.** Set an allowance per plan and enforce it in the product, not only on the pricing page. In Upwork Scout each user gets 150 AI scores a day. When the cap is hit, matching does not stop; it degrades to the deterministic filters for the rest of the day. The user still gets alerts, I still get a predictable bill.\n\n**Per agent run.** Any agent that can call a model repeatedly needs a ceiling on steps, tokens and wall-clock time for one task, after which it stops and reports where it got to:\n\n``` js\n// illustrative: one budget object per agent run\nconst budget = { maxSteps: 12, maxTokens: 60_000, deadlineMs: Date.now() + 90_000 };\n\nfunction canContinue(steps: number, tokensSoFar: number) {\n  return steps < budget.maxSteps\n    && tokensSoFar < budget.maxTokens\n    && Date.now() < budget.deadlineMs;\n}\n```\n\nAn agent that cannot do damage can still run up a bill. This is the cost twin of the containment rules I wrote about in [Give an AI Agent Write Access One Verb at a Time](https://dev.to/nabeelbaghoor/give-an-ai-agent-write-access-one-verb-at-a-time-32o6).\n\n**Per feature and per tenant.** Budgets with alerts, both at the provider and in your own logs, so a spike in one feature or one account is noticed in hours, not at month end. Decide in advance what happens when a budget is hit: fall back to a smaller model, queue the work, or switch the feature off with a clear message. A fallback decided calmly beats an outage decided under pressure.\n\n**Against abuse.** Public AI features attract people who want a free model. Rate limits, sign-in for expensive actions and input size checks stop one bad actor from becoming your biggest customer.\n\nSometimes a feature is well built and still costs too much for what it earns. Then the fix is commercial, not technical: charge for it separately, put it on higher plans, add an allowance with paid top-ups, or narrow it to the cases where it creates the most value. Better to make that call before launch, from a cost model built on realistic usage, than after users expect it for free.\n\nThe opposite happens too. Teams cut a valuable feature because the bill looks big in isolation, when the cost per unit is tiny against what the unit earns. Cost per unit of value settles that argument as well.\n\nIf I had to compress it into a checklist for a PR review:\n\nDo that and a growing AI bill means a growing product, not a growing problem.\n\nI originally wrote a longer, client-facing version of this for the [Null Studio blog](https://nullstud.io/blog/ai-cost-control/). Curious what others are doing here: do you enforce per-user limits in code, or still rely on provider-level alerts?", "url": "https://wpnews.pro/news/model-calls-are-metered-most-ai-features-are-built-like-they-are-free", "canonical_source": "https://dev.to/nabeelbaghoor/model-calls-are-metered-most-ai-features-are-built-like-they-are-free-1e4", "published_at": "2026-10-06 21:34:15+00:00", "updated_at": "2026-10-06 21:47:50.538820+00:00", "lang": "en", "topics": ["ai-agents", "ai-infrastructure", "large-language-models", "ai-products", "mlops"], "entities": ["CallGuard AI", "CallSetter AI"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/model-calls-are-metered-most-ai-features-are-built-like-they-are-free", "markdown": "https://wpnews.pro/news/model-calls-are-metered-most-ai-features-are-built-like-they-are-free.md", "text": "https://wpnews.pro/news/model-calls-are-metered-most-ai-features-are-built-like-they-are-free.txt", "jsonld": "https://wpnews.pro/news/model-calls-are-metered-most-ai-features-are-built-like-they-are-free.jsonld"}}