{"slug": "ai-agent-cost-per-month-the-token-maths-worked-out", "title": "AI Agent Cost per Month: The Token Maths, Worked Out", "summary": "A developer has published a token-cost formula for running tool-using AI agents, showing that a four-message support chat can consume over 45,000 input tokens — roughly twelve times Anthropic's own estimate of about 3,700 tokens per support conversation — because each turn resends the full prior history. The writeup derives history growth as k × h × T(T−1)/2, meaning quadrupling turns from 4 to 8 multiplies history tokens by about 4.7, and prices a worked small-business example at $170.94 per month on Claude Haiku 4.5. It also notes that Anthropic's newer tokenizer produces approximately 30% more tokens for the same text, so counts should be measured on the target model.", "body_md": "A four-message support chat with a tool-using AI agent can send the model more than 45,000 input tokens. That's about twelve times Anthropic's own estimate of roughly 3,700 tokens for a support conversation ([Anthropic pricing](https://platform.claude.com/docs/en/about-claude/pricing)). Nothing is broken: each turn simply resends everything that came before it.\n\nThat's why most guides to **AI agent cost per month** give you wide ranges instead of a number. This one gives you a formula, worked through a small-business example and priced against six current models using each provider's official pricing page.\n\nMost articles on this topic mix two different bills:\n\nOne top-ranking article says plainly that \"the published per-token price tells you almost nothing about what an agent will actually spend\" ([CloudZero](https://www.cloudzero.com/blog/ai-agent-cost/)). True, but behaviour can be measured — and anything measurable belongs in a formula.\n\nNot sure you need an agent at all? Read [Zapier, n8n or an AI agent](https://banxal.com/blog/zapier-n8n-or-ai-agent) — often one AI step inside a normal workflow is enough.\n\nFour things set your monthly bill:\n\nOne more catch: **the same text produces different token counts on different models.** Anthropic says the tokenizer in its newer models \"produces approximately 30% more tokens for the same text\" ([Anthropic pricing](https://platform.claude.com/docs/en/about-claude/pricing)). Count tokens on the model you'll actually use.\n\nMonthly run cost is conversations times per-conversation token cost, plus retry overhead, plus fixed costs. Here it is in full:\n\n```\nMonthly run cost = N × (I × p_in + O × p_out) / 1,000,000 × (1 + r) + fixed costs\n\nN      = conversations per month\nI      = input tokens per conversation (summed over every model call)\nO      = output tokens per conversation\np_in   = price per million input tokens; p_out = price per million output tokens\nr      = retry/loop overhead (0.10 = 10% extra calls)\nfixed  = vector DB, monitoring, hosting, search fees, human fallback\n```\n\nThe hard part is calculating I correctly. Each call's input is **fixed prefix + history so far + new content**. With T turns, k calls per turn and h tokens added to the history each turn:\n\n```\nHistory tokens per conversation = k × h × T(T − 1) / 2\n```\n\nThe T(T − 1) term is the important part. Going from 4 turns to 8 turns multiplies history tokens by 4.7 (from 6 to 28 in the T(T − 1)/2 term), not by 2.\n\nHere's the same calculation as a small pricing calculator you can edit:\n\n``` python\ndef agent_cost(convs, turns, prefix, user, tool_call, tool_result, answer,\n               price_in, price_out, overhead=0.10):\n    history, tokens_in, tokens_out = 0, 0, 0\n    for _ in range(turns):\n        tokens_in += prefix + history + user                            # call 1: pick a tool\n        tokens_in += prefix + history + user + tool_call + tool_result  # call 2: answer\n        tokens_out += tool_call + answer\n        history += user + tool_call + tool_result + answer\n    per_conv = (tokens_in * price_in + tokens_out * price_out) / 1e6\n    return per_conv * (1 + overhead) * convs\n\nprint(agent_cost(3000, 4, 2000, 100, 50, 1500, 250, 1.00, 5.00))  # ≈ 170.94\n```\n\nThis example prices a small-business support bot at $170.94 a month on Claude Haiku 4.5, with every step shown so you can swap in your own numbers.\n\n**The setup (these are assumptions, so swap in your own):** a support agent that searches a knowledge base and answers customers.\n\n| Assumption | Value | \n|---|---|\n| Conversations per month (N) | 3,000 (about 100 a day) | \n| Customer messages per conversation (T) | 4 | \n| Model calls per turn (k) | 2 (choose tool → answer) | \n| System prompt + tool definitions | 2,000 tokens | \n| Customer message / tool call / retrieved docs / answer | 100 / 50 / 1,500 / 250 tokens | \n| Retry overhead (r) | 10% | \n\n**Step 1: tokens added to history per turn.** 100 + 50 + 1,500 + 250 = **1,900**. History at the start of turns 1–4 is 0, 1,900, 3,800 and 5,700 tokens.\n\n**Step 2: input per turn.**\n\n**Step 3: input per conversation.** 4 × 5,750 = 23,000. History adds 2 × (0 + 1,900 + 3,800 + 5,700) = 22,800. That matches the formula: 2 × 1,900 × 6 = 22,800. **Total I = 45,800 tokens.**\n\n**Step 4: output per conversation.** (50 + 250) × 4 = **O = 1,200 tokens.**\n\n**Step 5: monthly tokens, with 10% overhead.**\n\n**Step 6: price it.** On Claude Haiku 4.5 ($1 input / $5 output per million tokens): 151.14 × $1 + 3.96 × $5 = $151.14 + $19.80 = **$170.94 a month**, or about **$0.057 per conversation**. The Python above gives the same figure.\n\nNotice the ratio: the agent reads 38 tokens for every token it writes. **For tool-using agents, the input price matters far more than the output price.**\n\nAcross six current models, the same worked example costs $17 to $342 a month in model fees alone. *Prices as of 3 October 2026, taken from each provider's official pricing page. Re-check them before relying on this; they change often.* Monthly cost uses the worked example (151.14M input, 3.96M output) with no caching and no batching.\n\n| Model | Input / Output per 1M tokens | Cached input per 1M | Example monthly cost | \n|---|---|---|---|\n| OpenAI gpt-6-luna | $0.10 / $0.50 | $0.01 | **$17.09** | \n| Gemini 3.5 Flash-Lite | $0.30 / $2.50 | $0.03 + storage | **$55.24** | \n| Claude Haiku 4.5 | $1 / $5 | $0.10 | **$170.94** | \n| Gemini 3.5 Flash | $1.50 / $9 | $0.15 + storage | **$262.35** | \n| Claude Sonnet 5.5 | $2 / $10 | $0.20 | **$341.88** | \n| OpenAI gpt-6.1-sol | $2 / $10 | $0.10 | **$341.88** | \n\nSources: [Anthropic](https://platform.claude.com/docs/en/about-claude/pricing), [OpenAI](https://developers.openai.com/api/docs/pricing) (short-context rates, up to 272K input tokens), [Google](https://ai.google.dev/gemini-api/docs/pricing) (Gemini caching also charges $1 per million tokens per hour of storage).\n\n**Prompt caching** is the biggest discount you can control. On Anthropic, a cache read costs 0.1× the base input price and a 5-minute cache write costs 1.25× (on Sonnet 5.5, that's $0.20 and $2.50 per million). With automatic caching, each new request reads everything up to the previous message from cache ([Anthropic prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching)). In our example, the eight calls add up to 45,800 input tokens. Of these, 36,450 are cache reads and 9,350 are writes:\n\nTwo catches: the default cache lasts 5 minutes (a 1-hour cache costs 2× to write), so a quiet customer breaks the chain; and Haiku 4.5 only caches prompts of at least 4,096 tokens, so the early calls here wouldn't be cached ([Anthropic prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching)).\n\n**Batch APIs** take 50% off at Anthropic, OpenAI and Google, but only for jobs that can wait, like nightly ticket summaries. They're no use for live chat.\n\n**Watch for introductory prices.** Gemini 3.8 Flash is $0.75 / $3.75 until 31 December 2026, then $1.50 / $7.50 from 1 January 2027 ([Google](https://ai.google.dev/gemini-api/docs/pricing)). If you budget at the introductory price, your bill will double when it ends.\n\nThe model fee is only part of the cost of running an AI chatbot. The rest:\n\nFor comparison, Intercom's Fin charges $0.99 per resolved conversation ([Intercom](https://www.intercom.com/pricing)) — not like-for-like, since that price includes the product itself, but a useful ceiling.\n\nRunning an open model on a GPU you rent turns a per-token bill into a per-hour one. On Lambda, one NVIDIA A10 (24 GB) is $1.29/hour and one H100 PCIe is $3.29/hour ([Lambda](https://lambda.ai/pricing)). Running all month (about 730 hours), that's **about $942 and $2,402**.\n\nOur whole example costs $17–$342 a month through an API. **At small-business volume, self-hosting costs more, often several times more.** It starts to make sense when:\n\nTo find your break-even point, benchmark how many tokens per second the GPU handles on your workload, then add the engineering time for patching, scaling and uptime.\n\nPull a week of logs and check these:\n\n`cache_read_input_tokens`) show whether caching is working.\nIf you want this formula applied to your real logs, or an agent built with budget caps and caching from the start, that's part of our [AI integration & agents](https://banxal.com/services) work. [Send us your volumes](https://banxal.com/contact) and we'll tell you what it should cost to run — or if an agent isn't worth it at all.\n\nIn our worked example, 3,000 support conversations cost $17–$342 a month in model fees across six current models, before caching. On top of that come vector storage, monitoring, search and your team's time on escalations.\n\nEvery call resends the system prompt, tool definitions and the conversation so far. History cost grows with the square of the number of turns, and retries add more full-price calls.\n\nUsually. In our example it cut the Sonnet 5.5 bill by 59%. Check your model's minimum cacheable prompt length and how long the cache lasts.\n\nRarely at small-business volume. One rented A10 GPU running all month costs about $942, which is more than our example's highest API bill.\n\nWhat does your agent's input-to-output ratio look like, and which cost driver surprised you most? Tell us in the comments.\n\n*Originally published on [banxal.com](https://banxal.com/blog/ai-agent-cost-per-month).*", "url": "https://wpnews.pro/news/ai-agent-cost-per-month-the-token-maths-worked-out", "canonical_source": "https://dev.to/manveer-banxal/ai-agent-cost-per-month-the-token-maths-worked-out-4e35", "published_at": "2026-10-03 07:53:15+00:00", "updated_at": "2026-10-03 08:07:57.188064+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-infrastructure", "ai-tools"], "entities": ["Anthropic", "Claude Haiku 4.5", "CloudZero", "Zapier", "n8n"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ai-agent-cost-per-month-the-token-maths-worked-out", "markdown": "https://wpnews.pro/news/ai-agent-cost-per-month-the-token-maths-worked-out.md", "text": "https://wpnews.pro/news/ai-agent-cost-per-month-the-token-maths-worked-out.txt", "jsonld": "https://wpnews.pro/news/ai-agent-cost-per-month-the-token-maths-worked-out.jsonld"}}