Where your Claude Code tokens actually go. Output is 0.2%. A developer counted tokens across 173 Claude Code sessions and roughly 60,000 turns, finding that all model output accounted for just 0.2% of tokens (56.6M) while re-sent conversation context made up 97.5% (25.3B), a roughly 450-to-1 ratio. After applying Anthropic's published cache-read and output pricing, output still only represents 7-14% of the bill, with re-reading and re-caching the session accounting for 86-93%. The finding matches earlier measurements in a Claude Code GitHub issue and a dev.to post, but the author adds a per-turn cost breakdown showing that a token's true cost is its size multiplied by the number of turns it remains in context. You hit the usage limit before lunch. So you do what everyone online says: make the AI talk less. Ban the "Great question ", cut the summaries, install one of the skills that trims its replies. I wanted to know where the tokens actually go, so I counted. Not estimated. Counted, from the transcripts Claude Code already keeps on your machine: every session still on mine, 173 of them, about 60,000 turns. Everything Claude wrote back to me was 0.2% of the tokens. The rest was my own session, fed back in. That number is real, but it isn't the whole story either, because most of those re-sent tokens are cheap ones. This post is both halves: where the tokens go, what they really cost, and the few things that actually move the number. Here's the summary across every project, straight from the terminal: bash $ baggage --all baggage — 173 sessions, all projects, 60,454 turns before you typed anything 75k tokens system prompt, every tool schema, every skill, CLAUDE.md. paid again on all 60,454 turns = 4.5B tokens, 17% of the bill, whether you called any of it or not. conversation you can see 27.8M tokens what the API billed 26.0B tokens a 934x gap. Some of it is the fixed cost above; the rest is everything you picked up being re-sent on every later turn. output tokens 56.6M 0.2% re-sent context 25.3B 97.5% <- the bill Two lines matter. Output every reply, every line of code Claude wrote, its thinking included is 56.6 million tokens. Re-sent context is 25.3 billion. That's roughly 450 to 1. I'm not the first to see this shape. A GitHub issue https://github.com/anthropics/claude-code/issues/24147 on Claude Code found the same thing in 30 days of someone else's transcripts, and a dev.to post https://dev.to/ploofnexa/i-measured-where-claude-code-actually-spends-tokens-968-is-re-reading-history-my-typing-was-16gm measured 0.5% output across 32 sessions. Different people, different work, same picture. What none of them show is which things in a session cost the most, and what any of it means for the bill. That's the rest of this post. The model has no memory between turns. Every time you press enter, Claude Code sends the whole conversation again: the system prompt, the tool definitions, every file it read, every command output, every reply so far, and then your new message. Think of a suitcase you have to carry up every flight of stairs. Pack a brick on the second floor of a 200-floor building and you carry it up 198 more flights. A 4,000-token build log that lands at turn 12 of a 200-turn session isn't 4,000 tokens. It's 4,000 sent 188 more times. So the cost of anything in a session isn't its size. It's its size times the number of turns it stays. This is the part most guides get half right. Claude Code caches the conversation, and Anthropic prices a cache read https://platform.claude.com/docs/en/about-claude/pricing at a tenth of normal input a twentieth on Opus 5.5, a fortieth on Fable 5.1 . Output costs five times input. So 97.5% of the tokens is not 97.5% of the money. I priced my own split at those published rates, with the one-hour cache a subscription uses. Output comes to 7 to 14% of the cost , depending on the model. Re-reading and re-caching the session is the other 86 to 93% . Less dramatic than 0.2%. Still the opposite of what "make it talk less" assumes. On a Pro or Max plan you never see dollars, you see a limit. Anthropic's cost docs https://code.claude.com/docs/en/costs say re-read history still draws on your usage, at the cached rate, so a one-line question late in a long session draws usage for the whole session. How heavily a cached token counts against the limit isn't published. I've seen people online insist it's full weight, and others insist it's free. Nobody I found has shown either, so I'm not going to guess. That first block of the output is the one nobody shows you. Before you've typed a word, a session already carries the system prompt, the built-in tools, the descriptions of every skill you've installed, your CLAUDE.md files and anything else that loads at start. A typical session on my machine starts at about 75,000 tokens. It rides along on every turn. Across all my sessions, that floor alone is 17% of everything billed , whether I used any of it or not. It's also the one line you can fix in thirty seconds: every paragraph of CLAUDE.md you don't need is paid again on every turn of every session, and so is the description of every skill you never call. Here's one project, this website, with the list of what's costing the most rent, trimmed for length: bash $ baggage baggage — 13 sessions in singhlabs, 9,307 turns before you typed anything 84k tokens system prompt, every tool schema, every skill, CLAUDE.md. paid again on all 9,307 turns = 779.0M tokens, 18% of the bill, whether you called any of it or not. output tokens 7.1M 0.2% re-sent context 4.2B 98.0% <- the bill HEAVIEST THINGS YOU ARE STILL CARRYING rent = its size x the turns it stayed in context 304.4M 19.4% 3769x assistant reply 282.4M 18.0% 1376x your message 244.2M 15.6% 1415x claude-in-chrome · browser batch 54.6M 3.5% 145x WebSearch 25.6M 1.6% 275x claude-in-chrome · javascript tool 21.5M 1.4% 17x claude-in-chrome · get page text 14.9M 0.9% 159x Agent Two things surprised me. Claude's own replies are the biggest line: 19.4%. Not because they were expensive to write. Writing all of them was part of the 0.2%. They're expensive because each one stays in the session and is sent again on every turn after it. So the "talk less" skills aren't wrong. They're right for the wrong reason: a short reply saves you far more in re-sends than it ever cost to write. Browser automation is close behind: 15.6%. And that's the text alone: page contents, element lists, logs of each click. baggage doesn't count images, so every screenshot rides along on top of that figure, uncounted. A session that clicks through a website carries a stack of pictures of it. Ranked by what the counts above say, not by what's easiest to write about: Start fresh between unrelated tasks. /clear drops everything you've been carrying, and Anthropic's docs say it costs nothing. The brick stays on floor two. Lower the floor. Uninstall skills you don't use, keep CLAUDE.md to what Claude can't work out on its own, and run /context once to see what loads before you type. Keep big output out of the main session. Send a long log to a file and search it, instead of printing it into the chat where it's carried to the end. Hand noisy exploration and browser work to a subagent, so only its answer comes back. Don't break the cache mid-session. Per the caching docs https://code.claude.com/docs/en/prompt-caching , switching model, changing effort, turning on fast mode or changing MCP servers can throw it away, and then the whole session is written to cache again at the higher rate. Decide those at the start. Shorter replies, for the right reason. Ask for the answer without the essay. It helps, because the essay gets carried. On a paid plan, /usage now shows which skills, subagents and MCP servers used your allowance. Worth a look before changing anything. The tool that printed everything above is called baggage https://singhlabs.dev/baggage/ . It's free, it reads the transcripts already on your machine, and nothing leaves it. The name on npm belongs to someone else, so install it from the repo: bash $ npm install -g github:manpreet171/baggage $ baggage this project $ baggage --all every project What it doesn't tell you, so you don't over-read it: Tokens, not dollars. The totals are the exact counts the API reported. Pricing them is your model and your plan. Item sizes are estimates. The list ranks things with a rough four-characters-a-token ruler, the same ruler for everything. The totals above it are exact. Only what's still on your machine. Claude Code keeps about 30 days of history by default. One heavy user. These are my numbers. Yours will differ, which is the point of running it. This is how we work. Measure what's really happening before changing anything, then change the thing the numbers point at. It's the same rule we use when an AI system we build for a business starts costing more than it should. Sources: both terminal blocks are real baggage runs on my own machine on 5 Oct 2026, the second trimmed to whole lines for length. Prices and cache multipliers are from Anthropic's pricing page https://platform.claude.com/docs/en/about-claude/pricing ; the cost split weights my exact token counts by them. Cache and usage behaviour is from Claude Code's costs https://code.claude.com/docs/en/costs and prompt caching https://code.claude.com/docs/en/prompt-caching docs. Read next: The best CLAUDE.md rules are hiding in your chat history https://singhlabs.dev/blog/claude-md-rules-chat-history/ · I researched loop engineering to build a product. I built nothing. https://singhlabs.dev/blog/loop-engineering-map/ Originally published at singhlabs.dev https://singhlabs.dev/blog/claude-code-token-usage/ .