Why AI Providers Shouldn’t Grade Their Own Bills A developer argues that AI model providers should not be the sole authority on their own billing, proposing an independent "second ledger" that preserves the provider's usage report and checks it against delivery, completion, pricing, and the result the application actually expected. The piece notes that token metering can be accurate while a response is unusable — such as a truncated tool call missing a required account_id — forcing apps to pay twice for one job. It points to existing signals like OpenAI's incomplete status and incomplete reason fields and Google Gemini's separate candidate, thinking, and cached token metadata as useful but insufficient for determining whether a failure warrants a refund. Why the future of AI routing is an independent layer that credits back what providers break. Let’s look at a bill that looks completely normal. Your app asks a model to return a JSON object. The model starts well enough, gets halfway through a field, and stops. Your parser has to respond. Your app retries. The second attempt works, the user never sees the first one, but the dashboard does and it records two API calls. Somewhere in the usage column, the broken attempt sits quietly beside the useful one. There may be no billing error here. Tokens were processed and generated. The meter counted them. But you paid for a response your software could not use, then paid again to finish the job. That distinction is the point. A provider can bill correctly under its terms while delivering a failed result for your application. We need to account for both facts. Right now, the industry counts tokens precisely and broken work rather poorly. Compute has always charged for resources rather than outcomes. An EC2 instance bills while it is running, even if your application on it crashes or sits idle. That bargain is clear: you rent the machine; you own the code. AWS describes the charge in terms of the instance’s running state https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-instance-lifecycle.html . With a model API, you control much less of the execution path. The provider decides how to run the model, when to stop it, what usage to report, and how to price that usage. You buy input and output tokens. Your app needs an answer that meets its request. Those are different units. If you ask for prose, a cut off paragraph may be mildly annoying. If you ask for a tool call with a required account id, a response ending after “account ” is a failed transaction. The tokens on the invoice are real and so is the failure. The cause matters. Did the model hit an output limit you set? Did the provider end the stream? Did your client stop reading? Each calls for a different remedy, but a month-end invoice cannot tell you which happened. Engineering sees retries, finance sees spend, and support sees a user error. Those should be three views of the same event, tied to its charge. Provider billing teams are not sitting around plotting how to collect forty cents from your malformed JSON. That would be a terrible if true but no. The structural issue is simpler. The provider runs the service, reports usage, publishes the billing rules, and often holds the most detailed account of what happened inside its system. You get an API response, your own logs if you kept them, and are delivered an invoice later. When those records disagree, the provider’s record tends to be the starting point for the conversation. Even when everyone acts in good faith, customers need their own receipt. What was requested, what arrived, and what was charged. To be fair, providers do expose useful signals. OpenAI’s response events can include an incomplete status, an incomplete reason, and token usage details. Google’s Gemini API exposes separate usage metadata for candidate, thinking, and cached tokens. These are valuable fields, and they make independent checking possible in some cases. They do not, by themselves, tell you whether your application received a usable answer or whether a given failure qualifies for a refund. OpenAI’s streaming reference https://platform.openai.com/docs/api-reference/responses-streaming/response/refusal and Google’s token documentation https://ai.google.dev/gemini-api/docs/generate-content/tokens show how much nuance is already present in the raw records. We want customers to have a second ledger. It should preserve the provider’s usage report and check it against delivery, completion, pricing, and the result that the app expected. A modest, verifiable finding beats a dramatic refund estimate made from guesswork. There is an important distinction here. A billing error means the charge conflicts with measurable usage, a rate, a discount, or a published rule. A service failure means the call did not produce the result your application needed. A single event can be both. It can also be only one of them. The accounting should keep those possibilities separate. The failure modes are mundane, which makes them easy to bury in ordinary traffic. An empty response with billed usage. Zero visible characters alone does not prove the meter is wrong, models may use billed reasoning tokens, return a tool call, or represent a refusal separately. The question is whether the call delivered anything usable under the agreed contract after those token classes and output types are accounted for. A truncated response. The model returns most of a JSON object, then stops before the closing braces. If your output cap was too low, fix your configuration. If the upstream stream ended unexpectedly, investigate the route. Either way, the receipt needs the finish reason, bytes delivered, usage, and schema result to show what the failed call cost. A retry cascade. An agent hits an error, retries, times out, retries again, and finally succeeds. The customer sees one task; the provider may see several attempts. We would not claim every error is billed. We do want to know which attempts were processed and charged, and what the whole task cost. The successful final call is a flattering but incomplete number. A token or cache mismatch. Visible text is not always the full token count; hidden reasoning and other token classes can appear in usage. A suspected cache hit is not a confirmed discount, either. Compare the provider’s usage categories, applicable model rates, and observed charge before calling anything an overcharge. Sometimes the result is a strong billing candidate. Sometimes it is a question for review. Here is a small example. Record only attempt 3 and you see success. Record all 3 and you see what success cost. OpenRouter has already shown that a clearer remedy is possible for one narrow category: it announced zero charge for responses with zero completion tokens under its stated conditions https://openrouter.ai/blog/announcements/never-pay-for-empty-ai-responses-again/ . That is a useful precedent. It also illustrates why the details matter. Zero completion tokens is a specific rule; it does not automatically cover every partial stream, invalid schema, retry, or answer a customer cannot use. If you have seen a failure mode we missed, bring it to inferock-bench on GitHub https://github.com/inferock/inferock-bench . The measurement rules are public, and issues help us make the receipts more accurate. Many teams send traffic through multiple providers, frameworks, and fallback routes. Then finance asks why inference spend rose and gets four dashboards with four definitions of “request.” A routing layer gives selected traffic one accountable path and records the call while it happens. You cannot reconstruct a broken stream from a month-end invoice. You need the requested output, upstream route, bytes delivered, finish state, usage, and retries. A receipt created in the request path keeps those facts together and connects multiple attempts to the task that caused them. That is the idea behind Inferock. You send selected inference traffic through one accountable path; the path keeps route, timing, response, usage, and charge evidence attached to each observed call. It can report objective failures such as empty or malformed output when the evidence supports them, and it can keep weaker findings in a separate review category. The methodology https://inferock.ai/methodology/ describes the measurements and their limits. There are two ways to use that accountability. The difference matters. With Bring Your Own Keys BYOK , you keep your provider accounts and pay those provider bills directly. Inferock measures traffic and gives you receipts and loss reporting. If the evidence points to a provider billing issue, you have a better record to investigate or dispute it. With Managed inference , Inferock operates the upstream provider relationship for approved traffic. That creates room for a service credit when an eligible, objective failure meets the agreed terms. See current pricing and mode description https://inferock.ai/pricing/ . The rule starts with evidence. Did the call pass through the measured path? Is there a priced usage or charge record? What failed, and was the cause upstream, client-side, or a configuration limit? Does it meet the service terms? The receipt should show the answer and any gap in the evidence. There is a tempting version of this argument where every ugly model response is called an overcharge and every provider is cast as a villain. It would make for a loud post. It would also make it harder for anyone serious to build a fair system around the problem. Here is the standard we are willing to defend. Every material model call should have a customer-readable receipt. That receipt should connect the request, upstream route, delivered response, completion state, usage, applicable rate, observed charge, and any retry to one logical operation. It should say what can be verified and what cannot. It should separate a provider-recognized billing error from a service failure measured by an independent standard. We also want teams to stop accepting “the API returned 200” as proof that the job was done. A valid HTTP status is useful. A valid result for your customer is the thing you actually needed. The gap between them has a cost, and that cost should be measurable. We’re building Inferock around that conviction. You should be able to point to a call and say: this is what we asked for, this is what arrived, this is what the meter recorded, this is what failed, and this is what happens next. If the evidence is incomplete, we should say that in the same place. If a Managed failure qualifies for a credit, the credit should be visible there too. Providers will keep improving their meters. They should. We want them to publish clearer rules for partial streams, failed calls, hidden token classes, and cache charges. Their own credit policies should become easier to understand and use. But the only record of a failed purchase should not come from the company that sold it. That is the standard we want to see across AI infrastructure: a bill you can check, a failure you can prove, and a remedy whose limits are clear before you need it. If you want to see this on your own traffic, request an Inferock invite https://inferock.ai/ waitlist . If your model spend is growing, you already have enough uncertainty in the system. Your invoice should not be another one.