AI Agent Cost per Month: The Token Maths, Worked Out A developer has published a token-cost formula for running tool-using AI agents, showing that a four-message support chat can consume over 45,000 input tokens — roughly twelve times Anthropic's own estimate of about 3,700 tokens per support conversation — because each turn resends the full prior history. The writeup derives history growth as k × h × T(T−1)/2, meaning quadrupling turns from 4 to 8 multiplies history tokens by about 4.7, and prices a worked small-business example at $170.94 per month on Claude Haiku 4.5. It also notes that Anthropic's newer tokenizer produces approximately 30% more tokens for the same text, so counts should be measured on the target model. A four-message support chat with a tool-using AI agent can send the model more than 45,000 input tokens. That's about twelve times Anthropic's own estimate of roughly 3,700 tokens for a support conversation Anthropic pricing https://platform.claude.com/docs/en/about-claude/pricing . Nothing is broken: each turn simply resends everything that came before it. That's why most guides to AI agent cost per month give you wide ranges instead of a number. This one gives you a formula, worked through a small-business example and priced against six current models using each provider's official pricing page. Most articles on this topic mix two different bills: One top-ranking article says plainly that "the published per-token price tells you almost nothing about what an agent will actually spend" CloudZero https://www.cloudzero.com/blog/ai-agent-cost/ . True, but behaviour can be measured — and anything measurable belongs in a formula. Not sure you need an agent at all? Read Zapier, n8n or an AI agent https://banxal.com/blog/zapier-n8n-or-ai-agent — often one AI step inside a normal workflow is enough. Four things set your monthly bill: One more catch: the same text produces different token counts on different models. Anthropic says the tokenizer in its newer models "produces approximately 30% more tokens for the same text" Anthropic pricing https://platform.claude.com/docs/en/about-claude/pricing . Count tokens on the model you'll actually use. Monthly run cost is conversations times per-conversation token cost, plus retry overhead, plus fixed costs. Here it is in full: Monthly run cost = N × I × p in + O × p out / 1,000,000 × 1 + r + fixed costs N = conversations per month I = input tokens per conversation summed over every model call O = output tokens per conversation p in = price per million input tokens; p out = price per million output tokens r = retry/loop overhead 0.10 = 10% extra calls fixed = vector DB, monitoring, hosting, search fees, human fallback The hard part is calculating I correctly. Each call's input is fixed prefix + history so far + new content . With T turns, k calls per turn and h tokens added to the history each turn: History tokens per conversation = k × h × T T − 1 / 2 The T T − 1 term is the important part. Going from 4 turns to 8 turns multiplies history tokens by 4.7 from 6 to 28 in the T T − 1 /2 term , not by 2. Here's the same calculation as a small pricing calculator you can edit: python def agent cost convs, turns, prefix, user, tool call, tool result, answer, price in, price out, overhead=0.10 : history, tokens in, tokens out = 0, 0, 0 for in range turns : tokens in += prefix + history + user call 1: pick a tool tokens in += prefix + history + user + tool call + tool result call 2: answer tokens out += tool call + answer history += user + tool call + tool result + answer per conv = tokens in price in + tokens out price out / 1e6 return per conv 1 + overhead convs print agent cost 3000, 4, 2000, 100, 50, 1500, 250, 1.00, 5.00 ≈ 170.94 This example prices a small-business support bot at $170.94 a month on Claude Haiku 4.5, with every step shown so you can swap in your own numbers. The setup these are assumptions, so swap in your own : a support agent that searches a knowledge base and answers customers. | Assumption | Value | |---|---| | Conversations per month N | 3,000 about 100 a day | | Customer messages per conversation T | 4 | | Model calls per turn k | 2 choose tool → answer | | System prompt + tool definitions | 2,000 tokens | | Customer message / tool call / retrieved docs / answer | 100 / 50 / 1,500 / 250 tokens | | Retry overhead r | 10% | Step 1: tokens added to history per turn. 100 + 50 + 1,500 + 250 = 1,900 . History at the start of turns 1–4 is 0, 1,900, 3,800 and 5,700 tokens. Step 2: input per turn. Step 3: input per conversation. 4 × 5,750 = 23,000. History adds 2 × 0 + 1,900 + 3,800 + 5,700 = 22,800. That matches the formula: 2 × 1,900 × 6 = 22,800. Total I = 45,800 tokens. Step 4: output per conversation. 50 + 250 × 4 = O = 1,200 tokens. Step 5: monthly tokens, with 10% overhead. Step 6: price it. On Claude Haiku 4.5 $1 input / $5 output per million tokens : 151.14 × $1 + 3.96 × $5 = $151.14 + $19.80 = $170.94 a month , or about $0.057 per conversation . The Python above gives the same figure. Notice the ratio: the agent reads 38 tokens for every token it writes. For tool-using agents, the input price matters far more than the output price. Across six current models, the same worked example costs $17 to $342 a month in model fees alone. Prices as of 3 October 2026, taken from each provider's official pricing page. Re-check them before relying on this; they change often. Monthly cost uses the worked example 151.14M input, 3.96M output with no caching and no batching. | Model | Input / Output per 1M tokens | Cached input per 1M | Example monthly cost | |---|---|---|---| | OpenAI gpt-6-luna | $0.10 / $0.50 | $0.01 | $17.09 | | Gemini 3.5 Flash-Lite | $0.30 / $2.50 | $0.03 + storage | $55.24 | | Claude Haiku 4.5 | $1 / $5 | $0.10 | $170.94 | | Gemini 3.5 Flash | $1.50 / $9 | $0.15 + storage | $262.35 | | Claude Sonnet 5.5 | $2 / $10 | $0.20 | $341.88 | | OpenAI gpt-6.1-sol | $2 / $10 | $0.10 | $341.88 | Sources: Anthropic https://platform.claude.com/docs/en/about-claude/pricing , OpenAI https://developers.openai.com/api/docs/pricing short-context rates, up to 272K input tokens , Google https://ai.google.dev/gemini-api/docs/pricing Gemini caching also charges $1 per million tokens per hour of storage . Prompt caching is the biggest discount you can control. On Anthropic, a cache read costs 0.1× the base input price and a 5-minute cache write costs 1.25× on Sonnet 5.5, that's $0.20 and $2.50 per million . With automatic caching, each new request reads everything up to the previous message from cache Anthropic prompt caching https://platform.claude.com/docs/en/build-with-claude/prompt-caching . In our example, the eight calls add up to 45,800 input tokens. Of these, 36,450 are cache reads and 9,350 are writes: Two catches: the default cache lasts 5 minutes a 1-hour cache costs 2× to write , so a quiet customer breaks the chain; and Haiku 4.5 only caches prompts of at least 4,096 tokens, so the early calls here wouldn't be cached Anthropic prompt caching https://platform.claude.com/docs/en/build-with-claude/prompt-caching . Batch APIs take 50% off at Anthropic, OpenAI and Google, but only for jobs that can wait, like nightly ticket summaries. They're no use for live chat. Watch for introductory prices. Gemini 3.8 Flash is $0.75 / $3.75 until 31 December 2026, then $1.50 / $7.50 from 1 January 2027 Google https://ai.google.dev/gemini-api/docs/pricing . If you budget at the introductory price, your bill will double when it ends. The model fee is only part of the cost of running an AI chatbot. The rest: For comparison, Intercom's Fin charges $0.99 per resolved conversation Intercom https://www.intercom.com/pricing — not like-for-like, since that price includes the product itself, but a useful ceiling. Running an open model on a GPU you rent turns a per-token bill into a per-hour one. On Lambda, one NVIDIA A10 24 GB is $1.29/hour and one H100 PCIe is $3.29/hour Lambda https://lambda.ai/pricing . Running all month about 730 hours , that's about $942 and $2,402 . Our whole example costs $17–$342 a month through an API. At small-business volume, self-hosting costs more, often several times more. It starts to make sense when: To find your break-even point, benchmark how many tokens per second the GPU handles on your workload, then add the engineering time for patching, scaling and uptime. Pull a week of logs and check these: cache read input tokens show whether caching is working. If you want this formula applied to your real logs, or an agent built with budget caps and caching from the start, that's part of our AI integration & agents https://banxal.com/services work. Send us your volumes https://banxal.com/contact and we'll tell you what it should cost to run — or if an agent isn't worth it at all. In our worked example, 3,000 support conversations cost $17–$342 a month in model fees across six current models, before caching. On top of that come vector storage, monitoring, search and your team's time on escalations. Every call resends the system prompt, tool definitions and the conversation so far. History cost grows with the square of the number of turns, and retries add more full-price calls. Usually. In our example it cut the Sonnet 5.5 bill by 59%. Check your model's minimum cacheable prompt length and how long the cache lasts. Rarely at small-business volume. One rented A10 GPU running all month costs about $942, which is more than our example's highest API bill. What does your agent's input-to-output ratio look like, and which cost driver surprised you most? Tell us in the comments. Originally published on banxal.com https://banxal.com/blog/ai-agent-cost-per-month .