Drafted with AI help, human-reviewed by The Agent Loop.
Short version: You don't have a model bill. You have a loop bill with a model attached, and the loop is where the number comes from.
I ran the arithmetic on my own session while writing the cost post: 3,518,203 input tokens against 271,350 output tokens. Thirteen tokens read for every one written. The model wrote a word; it was handed a paragraph back, thirteen times over.
I have never once guessed this number right before checking. Every time I expect the output side to matter, and every time the input side has already decided the bill.
The best measurement I found is not a vendor slide. A study of about 4,300 real coding-agent sessions (Claude Code and Codex, 43 developers, roughly 350,000 LLM steps and 430,000 tool calls) reported the median step like this:
| Part of one LLM step | Median tokens |
|---|---|
| Cached prefix re-sent (system prompt, history, tools) | 119,000 |
| Newly appended text this step | 875 |
| Output the model writes | 214 |
That is per step, not per run; a run is dozens of these stitched together (source: arXiv 2606.30560, 2026-06). Read it as a ratio: that prefix is about 136Γ the text the step appends and roughly 550Γ the tokens the model writes.
Multiply by the steps. This is why a 10-step file-reading agent in a vendor guide came out at 472,500 input tokens versus 9,000 for a single pass, about 43Γ (Augment Code guide, illustrative, not an audit).
one step's anatomy (median, 4,300 sessions)
ββββββββββββββββββββββββββββββββββββββββββββββββ
β cached prefix re-sent every step 119,000 βββ your loop
ββββββββββββββββββββββββββββββββββββββββββββββββ€
β appended this step 875
β output written 214
ββββββββββββββββββββββββββββββββββββββββββββββββ
β² β²
0.1Γ if byte-identical you pay full
1.25Γ to write it once price for both
The mechanism is boring and fatal: the full conversation history (every prior prompt and completion) is carried forward unchanged on every round. Context accumulates, input grows superlinearly, and the study found that cache-read input tokens dominated both raw volume and dollar cost in every phase they analysed (arXiv 2604.22750, 2026-04).
The same trajectories show what humans do under that pressure: expensive runs re-open and re-edit the same files far more often, and the authors call it redundant back-and-forth that inflates context without proportional progress. The finding that matters: more tokens did not reliably produce higher accuracy.
Providers price the re-read differently from the fresh text, and the ratios are stable enough to plan around (Anthropic docs, 2026-09):
OpenAI's docs say caching is on by default, discounts run up to 90%, and the exact multiplier is model-dependent, so treat any single "OpenAI gives you 50% off" figure as the 2024 announcement, not the current table (OpenAI docs, 2026-09).
The practical reading: your prefix is either your cheapest token or your most repeated cost, and you choose which by whether it stays byte-identical. Change one header mid-run and the next step writes instead of reads.
On 500 SWE-bench Verified problems with eight frontier models run four times each, agentic coding tasks averaged roughly 3,500Γ the tokens of single-round code reasoning β and within the same task, runs varied up to 30Γ in tokens, with the worst run for a given problem about 2Γ the best (arXiv 2604.22750, 2026-04). The most expensive problem averaged about seven million more tokens than the cheapest.
So "what does a run cost" has no single answer. It has a distribution, and the width of that distribution is the thing worth managing.
Published per-task dollars exist, if you read the fine print. TheAgentCompany benchmarked OpenHands with Gemini 2.5 Pro at an average of $4.20 per task over 27.2 LLM-call steps at 30.3% success, with cost computed from token counts at API prices and no prompt caching assumed (arXiv 2412.14161, 2025-05). The same benchmark got Gemini 2.0 Flash under $1 per task at about 40 steps and lower performance.
Divide by the success rate and the honest number moves: $4.20 at 30.3% is roughly $13.90 per successful task β my division, not the paper's, and it assumes every failure's tokens were wasted. Failures often aren't free: Ο-bench found frontier function-calling agents succeeding on fewer than half of tasks, with pass^8 below 25% in the retail setting (arXiv 2406.12045, 2024-06).
Isn't 119K an unrealistic prefix? It is a median across real sessions with real tool schemas and long histories, not a toy example. If your prefixes are smaller, your loop is cheaper than the field, so measure instead of assuming.
Do caches make the input side free? No. Reads are cheap (0.1Γ) but they are billed on every step, and writes cost more than a normal token. Cheap is not free when it repeats a hundred times.
Why quote a 2024 benchmark? Because it is independent and methodological: token counts, step counts and success rates are all published. Vendor averages without n and method are the ones I throw away.
Should I switch models to cut cost? Try it after the loop: the same model on the same task already varied up to 30Γ in the benchmark above, so the loop is the bigger lever. Compare models once your prefix is stable and your successes are counted.
Do vendors make this worse? They price differently, not secretly. Vercel, for example, bills provider inference at cost plus $0.25 per million billable tokens: a stated pricing rule, not a market average (vendor docs, 2026-08). Whatever the markup, you still need per-run numbers to know what you are multiplying.
If this saved you a wrong guess about your own bill, tap the unicorn below β it takes one click and it is the only metric Dev.to actually shows me. And follow The Agent Loop if you want tomorrow's number in your feed: I am working through the arithmetic of running agents, one post a day.
Or get it by email instead: buttondown.com/theagentloop, one email per post and nothing else.