{"slug": "what-one-agent-run-actually-costs", "title": "What one agent run actually costs", "summary": "A developer's analysis of roughly 4,300 real coding-agent sessions (Claude Code and Codex, 43 developers, ~350,000 LLM steps) found that the median LLM step re-sends about 119,000 cached prefix tokens while appending only 875 new tokens and writing 214 output tokens, making the agent loop rather than the model the dominant cost driver. The same work reports that agentic coding tasks averaged roughly 3,500× the tokens of single-round code reasoning on 500 SWE-bench Verified problems, with run-to-run token variation up to 30× for the same task and no reliable accuracy gain from more tokens.", "body_md": "Drafted with AI help, human-reviewed by The Agent Loop.\n\n**Short version:** You don't have a model bill. You have a loop bill with a model attached, and the loop is where the number comes from.\n\nI ran the arithmetic on my own session while writing [the cost post](https://dev.to/theagentloop/your-agents-cost-problem-isnt-the-model-its-the-loop-13a3): **3,518,203 input tokens against 271,350 output tokens.** Thirteen tokens read for every one written. The model wrote a word; it was handed a paragraph back, thirteen times over.\n\nI have never once guessed this number right before checking. Every time I expect the output side to matter, and every time the input side has already decided the bill.\n\nThe best measurement I found is not a vendor slide. A study of **about 4,300 real coding-agent sessions** (Claude Code and Codex, 43 developers, roughly 350,000 LLM steps and 430,000 tool calls) reported the median step like this:\n\n| Part of one LLM step | Median tokens | \n|---|---|\n| Cached prefix re-sent (system prompt, history, tools) | **119,000** | \n| Newly appended text this step | **875** | \n| Output the model writes | **214** | \n\nThat is **per step**, not per run; a run is dozens of these stitched together (source: arXiv 2606.30560, 2026-06). Read it as a ratio: that prefix is about **136× the text the step appends** and roughly **550× the tokens the model writes**.\n\nMultiply by the steps. This is why a 10-step file-reading agent in a vendor guide came out at 472,500 input tokens versus 9,000 for a single pass, about **43×** (Augment Code guide, illustrative, not an audit).\n\n```\n  one step's anatomy (median, 4,300 sessions)\n  ┌──────────────────────────────────────────────┐\n  │ cached prefix re-sent every step  119,000 ◄── your loop\n  ├──────────────────────────────────────────────┤\n  │ appended this step                     875\n  │ output written                        214\n  └──────────────────────────────────────────────┘\n        ▲                          ▲\n   0.1× if byte-identical     you pay full\n   1.25× to write it once     price for both\n```\n\nThe mechanism is boring and fatal: **the full conversation history (every prior prompt and completion) is carried forward unchanged on every round.** Context accumulates, input grows superlinearly, and the study found that **cache-read input tokens dominated both raw volume and dollar cost in every phase they analysed** (arXiv 2604.22750, 2026-04).\n\nThe same trajectories show what humans do under that pressure: expensive runs re-open and re-edit the same files far more often, and the authors call it redundant back-and-forth that inflates context without proportional progress. The finding that matters: **more tokens did not reliably produce higher accuracy.**\n\nProviders price the re-read differently from the fresh text, and the ratios are stable enough to plan around (Anthropic docs, 2026-09):\n\nOpenAI's docs say caching is on by default, discounts run **up to 90%**, and the exact multiplier is **model-dependent**, so treat any single \"OpenAI gives you 50% off\" figure as the 2024 announcement, not the current table (OpenAI docs, 2026-09).\n\nThe practical reading: **your prefix is either your cheapest token or your most repeated cost, and you choose which by whether it stays byte-identical.** Change one header mid-run and the next step writes instead of reads.\n\nOn 500 SWE-bench Verified problems with eight frontier models run four times each, agentic coding tasks averaged roughly **3,500× the tokens of single-round code reasoning** — and within the same task, runs varied **up to 30×** in tokens, with the worst run for a given problem about 2× the best (arXiv 2604.22750, 2026-04). The most expensive problem averaged about seven million more tokens than the cheapest.\n\nSo \"what does a run cost\" has no single answer. It has a distribution, and the width of that distribution is the thing worth managing.\n\nPublished per-task dollars exist, if you read the fine print. **TheAgentCompany** benchmarked OpenHands with Gemini 2.5 Pro at an average of **$4.20 per task** over 27.2 LLM-call steps at **30.3% success**, with cost computed from token counts at API prices and **no prompt caching assumed** (arXiv 2412.14161, 2025-05). The same benchmark got Gemini 2.0 Flash **under $1 per task** at about 40 steps and lower performance.\n\nDivide by the success rate and the honest number moves: $4.20 at 30.3% is roughly **$13.90 per successful task** — my division, not the paper's, and it assumes every failure's tokens were wasted. Failures often aren't free: τ-bench found frontier function-calling agents succeeding on fewer than half of tasks, with pass^8 below 25% in the retail setting (arXiv 2406.12045, 2024-06).\n\n**Isn't 119K an unrealistic prefix?** It is a median across real sessions with real tool schemas and long histories, not a toy example. If your prefixes are smaller, your loop is cheaper than the field, so measure instead of assuming.\n\n**Do caches make the input side free?** No. Reads are cheap (0.1×) but they are billed on every step, and writes cost more than a normal token. Cheap is not free when it repeats a hundred times.\n\n**Why quote a 2024 benchmark?** Because it is independent and methodological: token counts, step counts and success rates are all published. Vendor averages without n and method are the ones I throw away.\n\n**Should I switch models to cut cost?** Try it after the loop: the same model on the same task already varied **up to 30×** in the benchmark above, so the loop is the bigger lever. Compare models once your prefix is stable and your successes are counted.\n\n**Do vendors make this worse?** They price differently, not secretly. Vercel, for example, bills provider inference at cost plus **$0.25 per million billable tokens**: a stated pricing rule, not a market average (vendor docs, 2026-08). Whatever the markup, you still need per-run numbers to know what you are multiplying.\n\nIf this saved you a wrong guess about your own bill, tap the **unicorn** below — it takes one click and it is the only metric Dev.to actually shows me. And **follow [The Agent Loop](https://dev.to/theagentloop)** if you want tomorrow's number in your feed: I am working through the arithmetic of running agents, one post a day.\n\nOr get it by email instead: [buttondown.com/theagentloop](https://buttondown.com/theagentloop), one email per post and nothing else.", "url": "https://wpnews.pro/news/what-one-agent-run-actually-costs", "canonical_source": "https://dev.to/theagentloop/what-one-agent-run-actually-costs-28i8", "published_at": "2026-09-28 15:46:43+00:00", "updated_at": "2026-09-28 15:51:11.601310+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-infrastructure", "mlops", "ai-tools"], "entities": ["Claude Code", "Codex", "OpenAI", "Anthropic", "SWE-bench Verified", "TheAgentCompany", "Augment Code", "The Agent Loop"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/what-one-agent-run-actually-costs", "markdown": "https://wpnews.pro/news/what-one-agent-run-actually-costs.md", "text": "https://wpnews.pro/news/what-one-agent-run-actually-costs.txt", "jsonld": "https://wpnews.pro/news/what-one-agent-run-actually-costs.jsonld"}}