Picture a document-triage agent your platform team built last spring. Support uses it to summarize tickets. Sales uses it to pull clauses out of contracts. Legal uses it to flag risky terms. All three call the same endpoint, and the endpoint calls the model provider with one API key. At the end of the month, the provider’s usage report shows one total for that key, and finance asks the obvious question: whose number is it? Nobody can answer, because the information needed to answer it was never recorded.
That is the core of the AI cost attribution problem. It looks like a reporting problem, but it is a data-capture problem that only becomes visible in reporting.
Menlo Ventures puts enterprise spending on generative AI at $37 billion in 2025, more than triple the $11.5 billion of a year earlier. The estimate comes from a bottom-up market model plus a survey of about 500 U.S. enterprise decision-makers.
Ownership has not kept pace. In DoiT’s AI Spend Reality Check, which Sapio Research ran in February 2026 among 500 finance leaders at U.S. and U.K. organizations with 1,000 or more employees, 79% reported AI-related cost overruns in the previous twelve months. Just 15% could work out AI ROI without significant bottlenecks. Among the obstacles, 36% cited unclear financial attribution, behind the pace of technological change (40%), and finance and engineering defining success differently (37%). Asked who is accountable for AI spend, 55% pointed to technology leadership and 53% to finance. Those figures add up to more than 100%, so in many organizations both functions claim the budget, which in practice often means neither owns it.
Chargeback and showback have allocated shared IT costs for decades. Showback reports each team’s consumption while the cost stays in a central budget; chargeback moves the cost onto the consuming team’s budget. Both work cleanly when a resource has one owner: a virtual machine, a database, a software license. Tag the resource with its owner and the allocation follows. Shared resources were always the hard part, handled with allocation keys such as headcount or usage share. AI agents turn the hard part into the normal case.
AI agents break the one-owner model in three ways.
The vocabulary of chargeback still applies. The tagging model underneath it does not.
The fix is to change the unit of attribution. Instead of asking who owns the agent, ask who caused each call. Identity has to travel with the request itself rather than live in an infrastructure tag.
In practice, every model call needs two identities, not one:
For a single-purpose service, executor and payer point to the same owner, which is why the problem stays hidden until the first shared agent ships. For a shared agent, they diverge on every call. I think of this as the two-identity rule: if a request record cannot name both the executor and the payer, you can't allocate that request's cost without guesswork. Beyond the two identities, a usable attribution record carries four more fields:
• a root trace ID shared by every call in a chain, so nested calls roll up to the request that started them;
• the use case or task type, so cost can be read per unit of work rather than as raw volume;
• the model, with input, cached input, and output tokens priced at the rate actually billed;
• the outcome: success, failure or retry, plus latency.
There are three places to record it, and each trades coverage against effort.
In application code. Each service attaches metadata to its own calls. It is flexible and free to start, but coverage depends on every team remembering to do it, in every service, indefinitely. The first agent written in a hurry breaks the ledger.
At the provider. OpenAI projects split usage by project, with project-scoped API keys and per-project spend limits. Anthropic workspaces offer workspace-scoped keys and monthly spend limits, with up to 100 workspaces per organization. Both are reliable and require no code. Their limits are scope and silos: each provider keeps its own ledger, a team using several providers has to stitch those ledgers together, and a project or workspace alone does not carry a payer identity through a shared agent.
At a gateway. A proxy between your applications and the providers issues its own credentials, usually called virtual keys, one per team, service, or agent, and logs every request with its key, model, tokens, cost, and outcome. Applications change a base URL and a key rather than their code. The result is one ledger across providers and complete capture of the traffic that crosses it. The trade-offs are real: the gateway adds a hop to the request path, must run with the same availability as everything else on that path, and traffic that bypasses it stays invisible.
For shared agents, the gateway pattern has one more advantage: identity can follow the caller instead of the agent. Support’s requests to the triage agent carry one identity, sales’ requests carry another. This requires one change inside the shared agent, which must forward the caller’s credential or an identifying header with each model call. After that, the payer is captured when the call happens, not reconstructed later. Most attribution schemes look correct in a demo and drift within a month. These are the usual reasons.
-
Retries. An agent that re-asks the model twice because the output failed validation pays for three full calls. Charge the cost to the payer, but report retry overhead as a separate line for the agent’s owner. The payer did not cause the failures; the agent’s owner can fix them.
-
Nested calls. Without a root trace ID propagated through the chain, an orchestrator’s work is billed to whichever sub-agent made the last model call. Propagate the original payer identity with the trace, the same way you propagate a request ID for logging.
-
Cached tokens. Major providers bill cached input at a different rate from standard input. If your ledger prices every token at the standard rate, attributed totals will not match the invoice, and the gap will be blamed on attribution itself. Record cached tokens separately and price them at the rate actually billed.
-
Shared fixed costs. Embedding a document corpus, maintaining a shared vector index, or running nightly evaluations are costs no single request causes. Decide explicitly whether to allocate them in proportion to usage or hold them centrally and show them separately. Either works. Leaving the decision open produces a different answer every month.
Unattributed traffic. Some calls will arrive without a usable identity: legacy scripts, a forgotten prototype, a vendor integration. Report this as its own figure, with its share of total spend. Spreading it silently across teams makes every team’s number wrong and hides the coverage gap you should be closing.
Attribution is trusted only when it adds up. Once a month, add attributed spend to unattributed spend and compare the total with the provider invoices. A small difference is normal because of timing and rounding. A large or growing one usually means one of two things: the price table is stale after a provider price change, or traffic is reaching providers without passing through your capture point. Both are fixable, and both stay invisible if nobody runs the check.
Attribution tells you who spent the money. It does not tell you whether the spend was worth it. A team that spends twice as much as another but resolves five times as many support tickets is not the problem, yet a spend-only report makes it look like one.
Once identity is recorded on every request, the next step becomes possible: join cost to outcomes your other systems already record, such as tickets resolved, pull requests merged or documents processed, and compare value per dollar rather than dollars alone. That comparison is where the useful conversations start, and you can't build it without per-request attribution underneath.
-
Start with showback. Show every team its numbers before anyone is billed. Disputes about the data are cheaper to settle before money moves.
-
Give every payer-and-agent pair its own identity, a credential or a forwarded header, with names a finance colleague can read.
-
Route all model traffic through one capture point, and measure what still bypasses it.
-
Give every credential a budget owner. A number with no person attached gets reported and ignored.
-
Set caps and alerts per credential, so a looping agent hits a limit within minutes rather than at month end.
-
Review weekly. By the time a monthly invoice arrives, the behavior behind a spike is weeks old.
-
Move to chargeback only once the numbers are trusted, and only if your accounting model needs it. Some organizations never do, and that is a legitimate choice.
AI spend is hard to attribute because the old unit of allocation, the resource, no longer maps to a single owner. Moving the unit to the request, recording executor and payer on every call and capturing both at one point solves most of the problem. Handling the edge cases (retries, nested calls, cached tokens, shared costs and unattributed traffic) solves the rest. None of it needs a new discipline. It needs one more fact recorded at the moment it is cheapest to record: when the call is made.
Disclosure: I work on tooling for AI agent governance, so this is a problem I think about professionally — and, yes, with some bias. I’ve kept this piece about the problem, not any product.
Chargeback Was Built for Servers. AI Agents Broke It was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.