Tokens | Context window | Caching| everything you got wrong… Developers routinely underestimate LLM costs because they misunderstand three mechanics: tokens are statistical text fragments rather than words, tool-definition JSON schemas are re-sent and re-billed on every request, and output tokens cost roughly five times input tokens because decode generates one token per forward pass while prefill parallelizes the whole sequence. The article advises developers to measure actual counts with Anthropic's client.messages.count_tokens API against model claude-opus-5 rather than relying on the "approximately four characters per token" rule, noting a single UUID can cost twelve to fourteen tokens and a dozen documented tools can consume four to five thousand tokens per request. Most developers discover the problem in production. The agent that ran fine during testing suddenly costs three times what they budgeted. Response times creep up as conversations go on. Something that worked on a short demo fails when a user asks four questions in a row. And the logs show nothing useful, because nothing actually crashed. The cause is almost always the same three things: a misunderstanding of what a token is, a misunderstanding of what the context window does, and zero thought given to caching. None of these are exotic failures. They are the default outcome when you build without understanding the machine underneath. This article goes after all three, in order, from the surface down. Most developers hear “token” and picture a word. One word, one token. That mental model is wrong enough to cause real problems. A token is a fragment of text, usually a few characters, that a model’s tokenizer carved out of your input during the training process. The carving is not based on grammar or meaning. It is statistical. Common sequences of characters that appear together often in the training corpus get assigned their own token. Rarer sequences get split up. So “the” is one token. “running” is one token. But “idempotency” might be three tokens: “idem,” “pot,” “ency.” Your API key? Every UUID looks different, so each character group has to be encoded individually. A single UUID can cost twelve to fourteen tokens. A list of ten UUIDs costs more than a short paragraph. Here is the practical consequence. Nobody writes agents that process only clean English prose. Agents process JSON. They process code. They process database records with IDs and timestamps and field names. All of that costs more per character than English prose does — sometimes dramatically more. The common advice is “approximately four characters per token.” That is an average over English prose. It does not hold for anything your agent actually works with. Run this through a tokenizer and count yourself: python from anthropic import Anthropicclient = Anthropic result = client.messages.count tokens model="claude-opus-5", system=your system prompt, tools=your tool definitions, messages=your conversation history, print f"Actual token count: {result.input tokens}" That number will be larger than you expect. It always is. Here is the piece that catches almost everyone. Your tool definitions ,the JSON schemas that describe what tools your agent can call, are sent to the model on every single request. Not once. Every request. A dozen well-documented tools can run to four or five thousand tokens. You pay for those tokens on turn one, turn two, turn fifteen, turn forty. For the entire life of every conversation. If your agent handles ten thousand conversations a month and each conversation averages eight turns, you are paying for those tool schemas eighty thousand times. Before you add your eighth tool with its detailed parameter descriptions, count what you are already sending: result = client.messages.count tokens model="claude-opus-5", tools=your tool definitions, messages= , empty — just the tools system="", print f"Tool definitions alone: {result.input tokens} tokens per request, forever" That number does not go away. It just multiplies. Look at any provider’s pricing page and output tokens cost roughly five times what input tokens cost. This is not a pricing choice. It reflects what the hardware is doing. When your prompt goes in the prefill phase, the GPU processes the entire sequence in one pass. All forty thousand tokens at once, in parallel. It is fast because parallelism is what GPUs are built for. When the model generates a response, the decode phase produces one token at a time. Each token requires a full forward pass through the model. Each pass depends on the token before it, so nothing can be parallelized. The process is memory-bandwidth-bound, which is the most expensive constraint modern hardware faces. PREFILL DECODEone pass over the entire prompt one pass per output token┌──────────────────────────────┐ ┌─┐ ┌─┐ ┌─┐ ┌─┐ ┌─┐ ┌─┐│ ████████████████████████████ │ ───▶ │█│▶│█│▶│█│▶│█│▶│█│▶│█│ ...└──────────────────────────────┘ └─┘ └─┘ └─┘ └─┘ └─┘ └─┘ 40,000 tokens sequential · unavoidable parallel · fast ~1 pass each Two things follow from this directly. Latency is controlled by output length, not input length. An agent that feels slow is almost always generating too much text. Doubling the system prompt adds maybe a hundred milliseconds. Doubling the response length doubles the wait. If you want your agent to feel faster, the answer is almost never a shorter prompt. It is a shorter answer. A verbose response costs twice. You pay output rates when the agent generates it. Then it goes into the conversation history, and you pay input rates again on every subsequent turn. A chatty response early in a long conversation gets billed a dozen times by the time the conversation ends. The context window looks like memory. You put the conversation history in. The agent refers back to earlier messages. Everything appears to have continuity. It is not memory. The model stores nothing between calls. Every time you call the API, you reconstruct the entire history and send it again. The model reads it fresh. What looks like memory is you paying to reconstruct it on every turn. This reframes the economics entirely. A ten-turn conversation is not ten requests. Each request includes every prior turn plus the new message plus every tool result plus the system prompt plus the tool definitions. The tenth request is not ten times more expensive than the first — it is closer to fifty times more expensive, because the total tokens sent grows with every turn. Turn 1 → sends 6,000 tokensTurn 5 → sends 22,000 tokensTurn 20 → sends 80,000 tokensTurn 40 → sends 160,000 tokens approaching window limits Total spend across a forty-turn conversation does not grow linearly. It grows quadratically. This is the curve that breaks agent economics in production. Break down a real request from the middle of a conversation and you get something like this: The user typed 40 tokens. You are paying for roughly 46,000. The user’s actual question is under a tenth of a percent of the bill. Everything else is architecture you chose. And since you chose it, you can measure it, change it, and cut it. They look similar from the outside but have different causes and require different fixes. The first is overflow. The window fills up, the request fails or the provider silently compacts the history, and you lose something the agent needed. This is the failure everyone prepares for and the least dangerous of the three, because it announces itself. You either get an error or an obviously confused response. Fixable. The second is dilution. Long contexts degrade attention unevenly. Information at the start and end of a long prompt is recalled reliably. Information in the middle is not. The effect grows with context length. A fact the agent quoted correctly at 8,000 tokens can be ignored at 80,000 — no error, no warning, a confident answer that is simply wrong. This is why “the model supports a million tokens” does not mean “you can put a million tokens in it and expect consistent behavior.” Capacity is not attention. The third is poisoning. The agent writes something incorrect into its own history — a hallucinated value, a misread tool result, an incorrect intermediate conclusion. That error is now in the context for every subsequent turn, and the agent treats its own prior output as established fact. Poisoning compounds. An agent twenty turns into a poisoned conversation is confidently reasoning from a premise that was never true, and the transcript reads as perfectly coherent. Overflow announces itself. Dilution is silent. Poisoning is a feedback loop that gets worse the longer it runs. For long conversations, implement a rolling summary. When the conversation history exceeds a threshold — say, 20,000 tokens — summarize everything older than the last three turns into a single paragraph and replace the raw turns with it. This flattens the growth curve dramatically. php def maybe compress history messages: list, threshold: int = 20 000 - list: from anthropic import Anthropic client = Anthropic result = client.messages.count tokens model="claude-opus-5", messages=messages, system="", if result.input tokens < threshold: return messages Summarize all but the last three turns to summarize = messages :-6 -6 because each turn is 2 messages user + assistant recent = messages -6: summary response = client.messages.create model="claude-opus-5", max tokens=512, messages= {"role": "user", "content": "Summarize the following conversation history in 3-5 sentences, " "preserving any specific facts, decisions, or values mentioned:\n\n" + "\n".join m "content" for m in to summarize if isinstance m "content" , str } , system="You summarize conversation history accurately and briefly.", summary message = { "role": "assistant", "content": f" Previous conversation summary : {summary response.content 0 .text}" } return summary message + recent For tool output, trim aggressively. A database query that returns a hundred records when the agent only needs the first three is wasting 97 records worth of tokens on every subsequent turn. Return exactly what the agent needs and nothing more. Keep verified facts in a typed object, not in the conversation transcript. If the agent confirmed an order ID, store it in state. Retrieve it directly when needed. Do not let the agent reconstruct it from what it said three turns ago. Every provider that offers prompt caching is offering the same underlying mechanism: if the start of your prompt matches the start of a prompt you sent recently, they skip processing the matching portion and charge you a fraction of the normal rate for it. The savings are real. Anthropic’s cache reads cost roughly 10% of normal input token rates. Cache writes cost about 125% of normal rates. Break-even on a short-lived cache entry is two requests. The mechanism is a prefix match on exact bytes. Not semantic similarity. Not “close enough.” Exact bytes. This one sentence explains why most caching setups fail silently. When you send a request to Anthropic, the content is rendered in a fixed order regardless of how you structured your API call: tool definitions → system prompt → conversation history Tool definitions go first, always. They sit at position zero in the token sequence. This has one consequence that most developers discover by accident: if your tool list changes between requests, the cache is invalidated for the entire request. System prompt, all of history, everything. A tool list that changes per user. A tool list built dynamically from user permissions. A tool list where one tool’s schema gets regenerated on each request with a slightly different property order. Any of these mean you pay full price on every request, for every user, every turn, and nothing in the API response will tell you it happened. The silent failures are more common than the obvious ones. A timestamp in the system prompt. If your system prompt contains datetime.now to tell the agent the current time, the cache entry is different every minute. You never get a cache hit. A UUID or request ID near the front. Same problem. Every request is unique by construction. Unsorted JSON serialization. json.dumps schema without sort keys=True produces key ordering that depends on dictionary insertion order, which Python does not guarantee across calls. Two calls with the same schema can produce different byte sequences and therefore miss the cache. A user’s name or ID interpolated into the system prompt. Every user has a different prompt, so no cache is shared across users. This misses the cache for every requestsystem = f"You are a helpful assistant. Today is {datetime.now }. User: {user.name}." This caches correctlysystem = "You are a helpful assistant." Pass dynamic values in the first user message instead, after the cached prefix You mark where the stable prefix ends using cache control: python from anthropic import Anthropicclient = Anthropic response = client.messages.create model="claude-opus-5", max tokens=1024, tools=TOOL DEFINITIONS, stable - same for all users, sorted schemas system= { "type": "text", "text": SYSTEM PROMPT, stable - no interpolation, no timestamps "cache control": {"type": "ephemeral"}, } , messages=conversation history, volatile - changes every turn That single marker on the last block of the system prompt caches the tool definitions and system prompt together. Everything after it — the conversation history — is volatile and does not need to be cached, because it changes on every turn anyway. There is a minimum prefix length before caching kicks in. On Claude models it ranges from 512 to 4,096 tokens depending on the specific model. A system prompt shorter than the minimum simply does not cache, and nothing in the response will tell you. If you want to confirm caching is working: print f"Cache writes: {response.usage.cache creation input tokens}" print f"Cache reads: {response.usage.cache read input tokens}" print f"Normal input: {response.usage.input tokens}" If cache read input tokens is zero across repeated requests that should share a prefix, an invalidator is running. Render two consecutive requests' prompts as strings, diff them, and look for anything that changed. The cause is almost always in the list above. Two failure modes hit agents specifically and not ordinary chat applications. Cache entries are found by searching backward through a bounded window of content blocks, roughly twenty blocks. An agent turn that fires six parallel tool calls appends twelve content blocks. Two turns like that and the cache entry from the start of the conversation is out of the lookback window. The entry exists — the match just cannot find it. Long tool-heavy turns need an intermediate cache breakpoint every ten or twelve blocks. The second trap is fan-out. When an orchestrator fires five parallel requests with the same system prompt, none of them can read a cache entry that another is still writing. The entry only becomes readable once the first response starts streaming. Fire all five at once and all five pay full price. Fire one, wait for its first streaming token, then release the rest. The cache is warm by the time they arrive. None of these are obscure edge cases. They are the default behavior of every agent that was built without accounting for them. The tool definitions sitting at position zero in every request, costing the same tokens every turn of every conversation across every user. The context growing quadratically as the conversation goes on, causing slowdowns that have nothing to do with the model’s capabilities. The cache that was supposed to cut costs but is silently invalidated by a timestamp in the system prompt or unsorted JSON keys or dynamic tool schemas. These compound. An agent with verbose responses, no history compression, and a broken cache does not cost two or three times what it should. It costs an order of magnitude more, gets slower as conversations lengthen, and produces outputs that quietly degrade in accuracy as the context fills up. All without throwing a single error. The fix is not exotic. Count your tokens before you assume. Measure the context at turn ten and at turn forty. Put datetime.now in the user turn, not the system prompt. Sort your JSON serialization. Check cache read input tokens and verify the cache is actually working. Most of the failures described here take an afternoon to fix. Finding them takes longer only because nobody knows to look. I’m currently open to a remote AI Engineering role , particularly around agentic systems, LLM applications, and production AI systems. If you’re building AI products and looking for an engineer who thinks beyond prompting — into architecture, reliability, performance, and cost, I’d be interested in connecting. Resume: \ resume link\ https://docs.google.com/document/d/1vpJlYGVLqD7Jxi2ZHOPWDRfGnZiujC4Z/edit?usp=sharing&ouid=118348581902294088450&rtpof=true&sd=true LinkedIn: \ LinkedIn\ www.linkedin.com/in/delight-bessie This is the first article in a series on building AI agents that work in production. Tokens | Context window | Caching| everything you got wrong… https://pub.towardsai.net/tokens-context-window-caching-everything-you-got-wrong-10b5a32f7456 was originally published in Towards AI https://pub.towardsai.net on Medium, where people are continuing the conversation by highlighting and responding to this story.