cd /news/ai-agents/why-your-agent-loop-costs-more-than-… · home › topics › ai-agents › article
[ARTICLE · art-148784] src=pub.towardsai.net ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Why Your Agent Loop Costs More Than Its Model Price Tag

Anthropic's published rate card lists Opus at $4 per million input tokens and Haiku at $1, but swapping an agent between the two models typically cuts a tool-heavy agent's bill by far less than the expected 75% because the agent loop, not the per-token price, drives cost, according to the article. Anthropic's own framing describes an agent as a loop in which the model plans a step, acts through a tool, reads the result, and decides whether to continue, with each stage billed separately and tool results re-sent as input tokens on every subsequent Messages API call. Batch processing of asynchronous, non-interactive requests runs at a flat 50% discount on both input and output tokens across all three model tiers, and on Opus 5.5 and newer prior turns' thinking blocks stay in context and are billed again as input tokens while Sonnet and Haiku strip them between turns.

by read7 min views4 publishedOct 10, 2026

Opus costs $4 per million input tokens. Haiku costs $1 [1]. Swap an agent from one to the other and you’d expect the bill to drop by roughly 75%. In a tool-heavy agent it often drops by a fraction of that, because the per-token price was never the biggest variable in the equation. The loop around the model is.

In this article:

Anthropic publishes a flat rate card: input tokens, output tokens, and two flavors of prompt-cache token, priced per model [1].

That table is where most cost conversations start and stop. Someone notices Opus output costs 4x what Haiku costs, suggests downgrading, and moves on. It’s a real lever, and the levers section below covers it properly, but it treats the agent like a single API call instead of what it actually is: a loop that calls the model, sometimes a tool, and then the model again, and again, until it decides it’s done. Every one of those calls carries its own input and output token count, and the ratio between them is set by your architecture, not your model choice.

Batch processing is the other number worth knowing before anything else: asynchronous, non-interactive requests run at a flat 50% discount on both input and output tokens, across all three model tiers [1]. If any part of your agent’s workload doesn’t need a response inside the same conversation turn, such as a background summarization or an overnight re-indexing job, that discount applies before any of the architectural levers below even come into play.

Anthropic’s own framing of agents is a loop: the model plans a step, takes an action through a tool, reads the result, and decides whether to continue or stop [2]. Each of those stages is a separate billed event, and they don’t cost the same amount.

A planning call bills the full system prompt and tool schema definitions as input tokens, on every single turn, unless that block is cached. A tool call itself is usually cheap to produce, since the model just needs to emit a structured call, but the result that comes back has to be read by the model on the next call, and it goes back in as input tokens, in full, every time the conversation continues. If a task needs six tool calls before the model is satisfied, that result doesn’t get billed once. It gets re-read, and re-billed, on each of the five calls that follow it, because the Messages API resends the full conversation history on every request.

Retries compound this in a way that’s easy to miss in a cost estimate. If a tool call comes back malformed, or a step needs to be redone, the agent doesn’t resend just the retry. It resends the system prompt, the tool schemas, and every prior tool result still in context, because that’s what a conversational turn requires. One retry on a long-running task can cost more than the original planning call that triggered it.

Extended thinking adds a less obvious wrinkle. Thinking tokens are billed as output tokens, tracked separately in the response under usage.output_tokens_details.thinking_tokens [3]. On Opus 5.5 and newer, prior turns' thinking blocks stay in context and get billed again as input tokens on the next call; on Sonnet and Haiku, they're stripped between turns and don't carry forward [3]. Two agents doing the same multi-step task on different model tiers can have meaningfully different cost curves for that reason alone, separate from the base per-token rate.

Here’s a representative shape for a single tool-using agent turn on Sonnet 5.5, built from the published rates rather than a specific logged bill, so you can see where the dollars concentrate rather than take one number as gospel.

In that shape, the planning call and the final synthesis call, the parts people actually think of as “the model doing the task,” are a small minority of the total. The tool-result re-injection and the retry, the parts that scale with how many steps the task takes and how often something goes wrong, dominate it. That’s the pattern worth internalizing: cost in an agent loop scales with the number of turns and the size of what gets carried forward on each one, not with how smart or expensive the underlying model is per token.

Anthropic’s own documentation makes the re-injection cost concrete. In a worked example for its context-editing feature, clearing stale tool results before the next call brought one conversation from 70,000 input tokens down to 25,000 on the same turn [4].

That’s a 64% reduction in input tokens for that turn, achieved without touching the model tier, the prompt, or the task itself. It’s the clearest evidence available that the loop, not the price sheet, is where the savings live.

Model tier selection still matters, and it’s the lever most teams reach for first; if your task doesn’t need Opus-level reasoning, Sonnet or Haiku at half to a quarter of the input cost is the simplest change available [1]. I went deeper on how to frame that tradeoff, beyond just picking the cheapest model that passes your evals, in Choosing Between Opus, Sonnet, and Haiku: A Cost Framework.

Prompt caching is the second lever, and usually the higher-leverage one for agents with a stable system prompt or tool schema set. A cache write costs 1.25x the base input rate for a 5-minute TTL or 2x for a 1-hour TTL, but a cache hit costs only 10% of the base input rate on most models, and as little as 5% on Opus [5]. For a system prompt and tool schema block that gets reused across dozens of turns in the same agent run, that discount compounds fast. I covered the caching mechanics in more depth, including the minimum cacheable token thresholds per model, in Prompt caching and token budgeting: what’s documented vs. exam-tested.

Context editing is the third lever, and the one most agent builders haven’t configured yet. The clear_tool_uses_20250919 edit type automatically clears the oldest tool results once a conversation crosses a configured input-token threshold, replacing them with a placeholder instead of carrying the full result forward indefinitely [4]. It's applied server-side, so the client doesn't need to maintain a separate trimmed copy of the conversation. This is the same mechanism behind the 64% reduction above, and it pairs with the kind of structured memory I covered in The Three Layers of Memory in the Claude Agent SDK, which separates short-term and long-term context.

Batch processing is the fourth lever, already mentioned above: a flat 50% discount for anything that doesn’t need a synchronous response [1].

Before implementing any of the above, the token-counting endpoint lets you estimate the cost of a specific prompt, system message, and tool set before you send it, and it’s free to call [6]. For a team trying to decide whether a planned agent architecture is affordable before building it, that’s a cheaper way to find out than running it and checking the invoice afterward.

None of these levers are worth implementing blind. The Messages API usage object reports input_tokens, output_tokens, cache_creation_input_tokens, and cache_read_input_tokens separately on every response, plus output_tokens_details.thinking_tokens when extended thinking is active [3]. Logging those four or five numbers per call, broken out by which step of the agent loop produced them, turns a vague complaint like "our Anthropic bill is higher than expected" into a specific, fixable finding, such as discovering that one retry path is responsible for most of a task's token count, instead of a guess.

That logging discipline overlaps heavily with what you’d want captured for debugging a failing agent in the first place, which I wrote about in What to Log Before Your Agent Fails in Production. If you’re already logging tool calls and their outcomes, adding token counts to the same log line costs almost nothing and tells you where the money went, not just where the agent went wrong.

The uncomfortable implication is that FinOps for agents looks less like a pricing negotiation and more like a profiling exercise. The model picker is the control everyone can see. The loop is where the spend actually happens, and it’s invisible until you log it.

If you enjoyed the article and wish to show your support, make sure to: Why Your Agent Loop Costs More Than Its Model Price Tag was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.

── more in #ai-agents 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/why-your-agent-loop-…] indexed:0 read:7min 2026-10-10 · —