# Your cost dashboard can't tell you which agent ran up the bill

> Source: <https://dev.to/theagentloop/your-cost-dashboard-cant-tell-you-which-agent-ran-up-the-bill-4eej>
> Published: 2026-09-28 05:25:11+00:00

Drafted with AI help, human-reviewed by The Agent Loop.

**Short version:** My own analytics dashboard said 52 views this morning while every per-post row printed `null`. The total was right, the attribution was broken, and the footnote even claimed the data didn't exist. That is what most agent cost tracking looks like: a correct number you can't decompose. A provider invoice tells you that spending rose and never which agent run, retry, or tool call raised it. OpenTelemetry's GenAI conventions define **no cost attribute in the published spec** (the proposal has been open since August 2026) and mark their token counters as "a proxy for cost approximation", and the span that wraps your tool call carries **no usage fields**. Price each span yourself, roll it up from the child inference spans, and alarm at runtime instead of waiting for the invoice.

**For skimmers**

`execute_tool` spans exist, Two days ago I wrote up [why an agent bill grows even when the model price doesn't move](https://dev.to/theagentloop/your-agents-cost-problem-isnt-the-model-its-the-loop-13a3). A reader, **[@hannune](https://dev.to/hannune)**, replied with the follow-up I deserved: “Tag cost per call is the one that took me too long to add” — and the question underneath it: how do you actually attribute spend to a span, not to the account?

Start with what attribution gives you today. Amazon's Cost Anomaly Detection uses Cost Explorer data, and the documentation states the delay plainly: **up to 24 hours**, with the detector running roughly three times a day. Cloud cost monitors fire when "today's total cost exceeds $10,000". Useful. Also, by construction, a day late.

Braintrust's cost guide puts the limitation better than I can: the invoice "can show that spending increased, but it cannot explain which customer, feature, prompt change, retry pattern, or **agent run** caused the increase."

The scale question is no longer academic. One widely reported month-long run of roughly a hundred agent instances billed **$1,305,088.81** across **603 billion** tokens and **7.6 million** requests, with a single day at $19,985.84 (reported figures, treat as un-audited). At that size, "the total went up" is not an explanation.

The scary part isn't the size of the bill. It's that you can't name the line that produced it.

Here is the step everyone skips. You count tokens, then you treat the count as cost. It isn't, because the same token bills at different rates depending on state.

`input_tokens` counts only tokens `total_input_tokens = cache_read + cache_creation + input_tokens`.
So "input tokens" in your log isn't a cost. It's a count that has to be split into three buckets and multiplied by three different rates, with a price table that changes per model and per date. Miss the split and you are off by an order of magnitude, in either direction.

The volume makes the split matter. In September 2025 Claude 4 Sonnet reportedly hit **100 billion tokens a day** on OpenRouter, of which **99% were input tokens** accumulated in a trajectory and **1%** was what the model generated. Nearly all of your spend lives in the tokens you keep re-sending, and those are exactly the tokens cache state re-prices.

This is about **tool- and run-level** granularity. Vendor dashboards attribute what they are able to price; the question is what happens at the tool call, where the span you need has no fields to carry it.

This is the part I did not expect when I went looking for standards support.

OpenTelemetry's GenAI semantic conventions define `gen_ai.usage.input_tokens` and `gen_ai.usage.output_tokens`, and the registry page marks them **Deprecated, moved to the GenAI repository**. In that repository every GenAI attribute still carries a **Development** badge, not Stable. Fine, that's how specs mature.

What's not fine: the token-metrics document says the counters "serve as a **proxy for cost approximation**", and there's **no cost, USD, or price attribute in the published registry**. I grepped the whole conventions repo: nine mentions of "cost", every one of them prose, zero attributes to set.

There *is* a proposal. **PR #443, "Add per-operation cost conventions (`gen_ai.usage.cost.*`)", has been open since 9 August 2026**, with issue #287 asking for the same thing since May 2025. The standard knows about the gap and has not closed it. Meanwhile two vendors already emit `gen_ai.usage.cost` on their own — and are filing issues against themselves because the value ships as a bare number with **no currency field** ([OpenLIT](https://github.com/openlit/openlit/issues/1666), [Laminer](https://github.com/lmnr-ai/lmnr/issues/2447)). A cost with no unit is half an attribute.

Worse for anyone hoping to attribute per tool: the `execute_tool {gen_ai.tool.name}` span exists, is typed INTERNAL, and its attribute table carries `gen_ai.tool.*` fields plus `error.type`. **No usage, no cost.** The span that says *this tool ran* cannot say *this tool cost*.

The spec's own discussion explains why nobody has bolted one on: tool and retrieval charges "are not model cost to begin with", so they were left out of the cost work entirely. Fair, as far as model billing goes. It also means that for the line item your CFO actually asks about — *which tool cost money* — there is no standard home for the number, only child inference spans you can roll up yourself.

The commercial tools are honest about the same gap:

`generation` and `embedding` observations track cost; other observation types carry none. Tool and agent spans are unpriced by design.
Add one more trap from the spec itself: a retried request's span SHOULD cover the logical operation "with all retries", which means SDK-level retries (the Python SDK retries certain errors **twice by default**) are folded into one span and invisible in the trace. Three billed attempts, one visible operation.

Nobody's standard says what a call cost. Everyone's dashboard prints a number anyway.

**Belief 1: the invoice is attribution.** It is a ledger, not a trace. AWS gives you up to 24 hours of lag and no run identity. You will always be able to say *that* week was expensive and never *which* agent.

**Belief 2: our observability tool tracks cost, so we're attributed.** It tracks cost on the spans it can price. Generation spans, maybe embeddings. The tool-call span that actually explains the spend is unpriced, so your per-feature rollup silently under-counts and your total silently over-counts. Datadog calls the result PARTIAL COST for a reason.

**Belief 3: the spend limit protects us.** OpenAI's spend limits and alerts are enforced at organization or project scope. Nothing there stops one runaway loop inside one project from being the whole bill. A budget at the account level is a smoke alarm in the building's lobby.

``` php
tool call
  |
  v
inference spans  -->  split billed tokens:  cache read | cache write | plain input
  |                                            |
  v                                            v
price table  <-----  keyed by model + date (not a constant)
  |
  v
price attributes on the PARENT span  -->  rollup: tool -> run -> feature
  |
  v
runtime budget alarm (seconds, not 24h)

invoice  -->  reconciliation only, never the source of truth
```

`cache_read`, `cache_creation`, and `cost.usd` (your own attribute, until the spec grows one) onto the tool span. That is the only place the tool's spend becomes queryable.`$3/M` is how you quietly drift 20% off reality. Store it as data next to the trace, not in code.`cost_tags` pattern: team, feature, customer, run id. The invoice will never carry these.
**The payoff is measurable.** Inference-time trajectory trimming (AgentDiet) strips the useless, redundant and expired context agents keep re-sending: across two LLMs and two benchmarks it cut **input tokens 39.9–59.7%** and **total cost 21.1–35.9%** while [agent performance stayed the same](https://arxiv.org/abs/2509.23586). You cannot see that win on an invoice. You can see it the moment cost lives on the span.

Admitted limit: I have not run this on a production bill. What I did run is smaller and embarrassing: this morning's dashboard printed a correct 52 with a broken per-post column, because one field name was wrong and a footnote asserted the data was unavailable. Totals that nobody can decompose will always lie to you eventually.

**How do you attribute LLM cost to a specific tool or span?**

Count billable tokens on each inference span, split into cache read, cache write, and plain input, multiply each by a price looked up for that model and date, then write the sum onto the parent `execute_tool` span and roll up by tool, run, and feature. No standard does this for you.

**Does OpenTelemetry have a cost attribute for GenAI?**

Not yet. The published conventions define token counters and describe them as "a proxy for cost approximation", with no cost, USD, or price attribute, and `execute_tool` spans carry no usage fields either. A proposal (`gen_ai.usage.cost.*`, PR #443) has been open since August 2026 and is still open, so treat any `gen_ai.usage.cost` you see in a vendor's output as a private extension with no agreed unit.

**Why does my observability tool show a different total than the invoice?**

Three usual causes: overlapping usage buckets double-counting, spans that carry no price data being silently skipped (Datadog reports this as PARTIAL COST), and usage counts that never included cache writes: OpenAI states its Agents API usage "cannot determine the exact model charge".

**What's the difference between tokens used and tokens billed?**

Used is what the model saw. Billed is what you pay for: cached input re-reads at 0.1×, fresh cache writes at 1.25×, and plain input at 1×. Anthropic's `input_tokens` counts only tokens after the last cache breakpoint, so the two numbers are not comparable.

**How do you alert on agent spend before the invoice arrives?**

Put a budget on the unit that can be stopped (per run or per loop), evaluated against your attributed per-span cost. Account-level spend limits and cloud anomaly detectors work as backstops, but AWS's detector itself has up to a 24-hour delay.

**Over to you:** if one tool call in your agent tripled its spend last week, what would you check first, and would that check actually name the tool? Reply below, I read every one.

And if this saved you from shipping another unattributed total, tap the **unicorn** (or like) and follow The Agent Loop — I read every reply, and the next one is about the thing nobody wants to audit.

Prefer an inbox to a feed? Every post also goes out by email: [subscribe at buttondown.com/theagentloop](https://buttondown.com/theagentloop).
