# What one agent run actually costs

> Source: <https://dev.to/theagentloop/what-one-agent-run-actually-costs-28i8>
> Published: 2026-09-28 15:46:43+00:00

Drafted with AI help, human-reviewed by The Agent Loop.

**Short version:** You don't have a model bill. You have a loop bill with a model attached, and the loop is where the number comes from.

I ran the arithmetic on my own session while writing [the cost post](https://dev.to/theagentloop/your-agents-cost-problem-isnt-the-model-its-the-loop-13a3): **3,518,203 input tokens against 271,350 output tokens.** Thirteen tokens read for every one written. The model wrote a word; it was handed a paragraph back, thirteen times over.

I have never once guessed this number right before checking. Every time I expect the output side to matter, and every time the input side has already decided the bill.

The best measurement I found is not a vendor slide. A study of **about 4,300 real coding-agent sessions** (Claude Code and Codex, 43 developers, roughly 350,000 LLM steps and 430,000 tool calls) reported the median step like this:

| Part of one LLM step | Median tokens | 
|---|---|
| Cached prefix re-sent (system prompt, history, tools) | **119,000** | 
| Newly appended text this step | **875** | 
| Output the model writes | **214** | 

That is **per step**, not per run; a run is dozens of these stitched together (source: arXiv 2606.30560, 2026-06). Read it as a ratio: that prefix is about **136× the text the step appends** and roughly **550× the tokens the model writes**.

Multiply by the steps. This is why a 10-step file-reading agent in a vendor guide came out at 472,500 input tokens versus 9,000 for a single pass, about **43×** (Augment Code guide, illustrative, not an audit).

```
  one step's anatomy (median, 4,300 sessions)
  ┌──────────────────────────────────────────────┐
  │ cached prefix re-sent every step  119,000 ◄── your loop
  ├──────────────────────────────────────────────┤
  │ appended this step                     875
  │ output written                        214
  └──────────────────────────────────────────────┘
        ▲                          ▲
   0.1× if byte-identical     you pay full
   1.25× to write it once     price for both
```

The mechanism is boring and fatal: **the full conversation history (every prior prompt and completion) is carried forward unchanged on every round.** Context accumulates, input grows superlinearly, and the study found that **cache-read input tokens dominated both raw volume and dollar cost in every phase they analysed** (arXiv 2604.22750, 2026-04).

The same trajectories show what humans do under that pressure: expensive runs re-open and re-edit the same files far more often, and the authors call it redundant back-and-forth that inflates context without proportional progress. The finding that matters: **more tokens did not reliably produce higher accuracy.**

Providers price the re-read differently from the fresh text, and the ratios are stable enough to plan around (Anthropic docs, 2026-09):

OpenAI's docs say caching is on by default, discounts run **up to 90%**, and the exact multiplier is **model-dependent**, so treat any single "OpenAI gives you 50% off" figure as the 2024 announcement, not the current table (OpenAI docs, 2026-09).

The practical reading: **your prefix is either your cheapest token or your most repeated cost, and you choose which by whether it stays byte-identical.** Change one header mid-run and the next step writes instead of reads.

On 500 SWE-bench Verified problems with eight frontier models run four times each, agentic coding tasks averaged roughly **3,500× the tokens of single-round code reasoning** — and within the same task, runs varied **up to 30×** in tokens, with the worst run for a given problem about 2× the best (arXiv 2604.22750, 2026-04). The most expensive problem averaged about seven million more tokens than the cheapest.

So "what does a run cost" has no single answer. It has a distribution, and the width of that distribution is the thing worth managing.

Published per-task dollars exist, if you read the fine print. **TheAgentCompany** benchmarked OpenHands with Gemini 2.5 Pro at an average of **$4.20 per task** over 27.2 LLM-call steps at **30.3% success**, with cost computed from token counts at API prices and **no prompt caching assumed** (arXiv 2412.14161, 2025-05). The same benchmark got Gemini 2.0 Flash **under $1 per task** at about 40 steps and lower performance.

Divide by the success rate and the honest number moves: $4.20 at 30.3% is roughly **$13.90 per successful task** — my division, not the paper's, and it assumes every failure's tokens were wasted. Failures often aren't free: τ-bench found frontier function-calling agents succeeding on fewer than half of tasks, with pass^8 below 25% in the retail setting (arXiv 2406.12045, 2024-06).

**Isn't 119K an unrealistic prefix?** It is a median across real sessions with real tool schemas and long histories, not a toy example. If your prefixes are smaller, your loop is cheaper than the field, so measure instead of assuming.

**Do caches make the input side free?** No. Reads are cheap (0.1×) but they are billed on every step, and writes cost more than a normal token. Cheap is not free when it repeats a hundred times.

**Why quote a 2024 benchmark?** Because it is independent and methodological: token counts, step counts and success rates are all published. Vendor averages without n and method are the ones I throw away.

**Should I switch models to cut cost?** Try it after the loop: the same model on the same task already varied **up to 30×** in the benchmark above, so the loop is the bigger lever. Compare models once your prefix is stable and your successes are counted.

**Do vendors make this worse?** They price differently, not secretly. Vercel, for example, bills provider inference at cost plus **$0.25 per million billable tokens**: a stated pricing rule, not a market average (vendor docs, 2026-08). Whatever the markup, you still need per-run numbers to know what you are multiplying.

If this saved you a wrong guess about your own bill, tap the **unicorn** below — it takes one click and it is the only metric Dev.to actually shows me. And **follow [The Agent Loop](https://dev.to/theagentloop)** if you want tomorrow's number in your feed: I am working through the arithmetic of running agents, one post a day.

Or get it by email instead: [buttondown.com/theagentloop](https://buttondown.com/theagentloop), one email per post and nothing else.
