Everyone talks about AI costs in aggregate. Dashboards show monthly spend, per-model totals, team budgets. Nobody shows you a receipt.
So here's one. A single AI agent task, priced line by line — every token, every tool call, every retry. This is a worked example from production-style workloads I describe in my reference architecture work (full paper: https://doi.org/10.5281/zenodo.23178788). The numbers are representative of a mid-complexity support-triage agent running on AWS Bedrock. Your numbers will differ. The shape of the receipt won't.
A customer support triage agent: read an incoming ticket, pull the customer's history, check the knowledge base, decide whether to auto-resolve or escalate, and draft the response. One task, one receipt.
| # | Line item | Tokens in | Tokens out | Unit cost basis | Cost |
|---|---|---|---|---|---|
| 1 | System prompt (loaded once, cached) | 1,850 | — | cached input | $0.0011 |
| 2 | Ticket text + customer history | 2,300 | — | standard input | $0.0069 |
| 3 | Knowledge base retrieval (3 chunks) | 4,100 | — | standard input | $0.0123 |
| 4 | Reasoning: triage decision | — | 420 | standard output | $0.0050 |
| 5 | Tool call: CRM lookup | 180 | 90 | in/out | $0.0016 |
| 6 | Reasoning: draft response | — | 680 | standard output | $0.0082 |
| 7 | Retry: first draft failed tone eval, regenerated | 6,200 | 710 | in (cached) + out | $0.0121 |
| 8 | Embedding: ticket for memory store | 340 | — | embedding | $0.0000 |
| 9 | Final response + audit log write | — | 350 | standard output | $0.0042 |
| Total | ~15,000 | ~2,250 | ~$0.051 |
Wait — $0.051, not $0.23? The $0.23 figure from my paper is the fully-loaded cost: it includes the amortized cost of the eval harness, the retrieval infrastructure, and the human review queue time for escalations. The raw inference receipt is five cents. The governed cost is twenty-three cents. Both numbers matter, and most teams track neither.
Three things jump out:
1. The retry cost more than the original attempt. Line 7 — a failed tone evaluation forced a regeneration — cost $0.0121, more than lines 4+6 combined. Retries are the silent budget killer in agent systems. Every eval gate you add has a cost; every gate you skip has a bigger one. The receipt lets you price the tradeoff instead of guessing.
2. Context is the bulk of input spend. Lines 2+3 are 6,400 input tokens — 43% of the total input. That knowledge-base retrieval is doing real work, but it's also the first place to optimize: better chunking, semantic caching of frequent queries, and prompt compression all attack this line directly.
3. The system prompt is nearly free — because it's cached. Line 1 costs a tenth of a cent thanks to prompt caching. Without caching, that 1,850-token system prompt gets re-priced on every call in the chain. If your provider supports caching and you're not using it, you're donating money.
You can't govern what you can't itemize. Monthly dashboards tell you that you spent $40K. The receipt tells you why — which line items to attack, which eval gates earn their keep, which retries are worth preventing.
Three practices that fall out of this:
The next time someone asks what your AI agents cost, don't open the dashboard. Show them a receipt.
Karmendra Pandey is a Practice Architect in AI & ML at TEK Systems, building production agentic AI on AWS. His reference architecture for cost-governed AI agents is at https://doi.org/10.5281/zenodo.23178788, and the companion implementation at https://github.com/karmendra8386/agentic-ai-aws-reference.