cd /news/ai-tools/where-your-claude-code-tokens-actual… · home › topics › ai-tools › article
[ARTICLE · art-145353] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=· neutral

Where your Claude Code tokens actually go. Output is 0.2%.

A developer counted tokens across 173 Claude Code sessions and roughly 60,000 turns, finding that all model output accounted for just 0.2% of tokens (56.6M) while re-sent conversation context made up 97.5% (25.3B), a roughly 450-to-1 ratio. After applying Anthropic's published cache-read and output pricing, output still only represents 7-14% of the bill, with re-reading and re-caching the session accounting for 86-93%. The finding matches earlier measurements in a Claude Code GitHub issue and a dev.to post, but the author adds a per-turn cost breakdown showing that a token's true cost is its size multiplied by the number of turns it remains in context.

by read7 min views3 publishedOct 5, 2026

You hit the usage limit before lunch. So you do what everyone online says: make the AI talk less. Ban the "Great question!", cut the summaries, install one of the skills that trims its replies.

I wanted to know where the tokens actually go, so I counted. Not estimated. Counted, from the transcripts Claude Code already keeps on your machine: every session still on mine, 173 of them, about 60,000 turns.

Everything Claude wrote back to me was 0.2% of the tokens. The rest was my own session, fed back in.

That number is real, but it isn't the whole story either, because most of those re-sent tokens are cheap ones. This post is both halves: where the tokens go, what they really cost, and the few things that actually move the number.

Here's the summary across every project, straight from the terminal:

$ baggage --all
baggage — 173 sessions, all projects, 60,454 turns

  before you typed anything              75k  tokens
    system prompt, every tool schema, every skill, CLAUDE.md.
    paid again on all 60,454 turns = 4.5B tokens, 17% of the bill,
    whether you called any of it or not.

  conversation you can see             27.8M  tokens
  what the API billed                  26.0B  tokens

  a 934x gap. Some of it is the fixed cost above; the rest
  is everything you picked up being re-sent on every later turn.

  output tokens           56.6M   0.2%
  re-sent context         25.3B   97.5%   <- the bill

Two lines matter. Output (every reply, every line of code Claude wrote, its thinking included) is 56.6 million tokens. Re-sent context is 25.3 billion. That's roughly 450 to 1.

I'm not the first to see this shape. A GitHub issue on Claude Code found the same thing in 30 days of someone else's transcripts, and a dev.to post measured 0.5% output across 32 sessions. Different people, different work, same picture. What none of them show is which things in a session cost the most, and what any of it means for the bill. That's the rest of this post.

The model has no memory between turns. Every time you press enter, Claude Code sends the whole conversation again: the system prompt, the tool definitions, every file it read, every command output, every reply so far, and then your new message.

Think of a suitcase you have to carry up every flight of stairs. Pack a brick on the second floor of a 200-floor building and you carry it up 198 more flights. A 4,000-token build log that lands at turn 12 of a 200-turn session isn't 4,000 tokens. It's 4,000 sent 188 more times.

So the cost of anything in a session isn't its size. It's its size times the number of turns it stays.

This is the part most guides get half right. Claude Code caches the conversation, and Anthropic prices a cache read at a tenth of normal input (a twentieth on Opus 5.5, a fortieth on Fable 5.1). Output costs five times input. So 97.5% of the tokens is not 97.5% of the money.

I priced my own split at those published rates, with the one-hour cache a subscription uses. Output comes to 7 to 14% of the cost, depending on the model. Re-reading and re-caching the session is the other 86 to 93%.

Less dramatic than 0.2%. Still the opposite of what "make it talk less" assumes.

On a Pro or Max plan you never see dollars, you see a limit. Anthropic's cost docs say re-read history still draws on your usage, at the cached rate, so a one-line question late in a long session draws usage for the whole session. How heavily a cached token counts against the limit isn't published. I've seen people online insist it's full weight, and others insist it's free. Nobody I found has shown either, so I'm not going to guess.

That first block of the output is the one nobody shows you. Before you've typed a word, a session already carries the system prompt, the built-in tools, the descriptions of every skill you've installed, your CLAUDE.md files and anything else that loads at start. A typical session on my machine starts at about 75,000 tokens.

It rides along on every turn. Across all my sessions, that floor alone is 17% of everything billed, whether I used any of it or not. It's also the one line you can fix in thirty seconds: every paragraph of CLAUDE.md you don't need is paid again on every turn of every session, and so is the description of every skill you never call.

Here's one project, this website, with the list of what's costing the most rent, trimmed for length:

$ baggage
baggage — 13 sessions in singhlabs, 9,307 turns

  before you typed anything              84k  tokens
    system prompt, every tool schema, every skill, CLAUDE.md.
    paid again on all 9,307 turns = 779.0M tokens, 18% of the bill,
    whether you called any of it or not.

  output tokens            7.1M   0.2%
  re-sent context          4.2B   98.0%   <- the bill

  HEAVIEST THINGS YOU ARE STILL CARRYING
  (rent = its size x the turns it stayed in context)

   304.4M   19.4%  3769x  assistant reply
   282.4M   18.0%  1376x  your message
   244.2M   15.6%  1415x  claude-in-chrome · browser_batch
    54.6M    3.5%   145x  WebSearch
    25.6M    1.6%   275x  claude-in-chrome · javascript_tool
    21.5M    1.4%    17x  claude-in-chrome · get_page_text
    14.9M    0.9%   159x  Agent

Two things surprised me.

Claude's own replies are the biggest line: 19.4%. Not because they were expensive to write. Writing all of them was part of the 0.2%. They're expensive because each one stays in the session and is sent again on every turn after it. So the "talk less" skills aren't wrong. They're right for the wrong reason: a short reply saves you far more in re-sends than it ever cost to write.

Browser automation is close behind: 15.6%. And that's the text alone: page contents, element lists, logs of each click. baggage doesn't count images, so every screenshot rides along on top of that figure, uncounted. A session that clicks through a website carries a stack of pictures of it.

Ranked by what the counts above say, not by what's easiest to write about:

Start fresh between unrelated tasks. /clear drops everything you've been carrying, and Anthropic's docs say it costs nothing. The brick stays on floor two.

Lower the floor. Uninstall skills you don't use, keep CLAUDE.md to what Claude can't work out on its own, and run /context once to see what loads before you type.

Keep big output out of the main session. Send a long log to a file and search it, instead of printing it into the chat where it's carried to the end. Hand noisy exploration and browser work to a subagent, so only its answer comes back.

Don't break the cache mid-session. Per the caching docs, switching model, changing effort, turning on fast mode or changing MCP servers can throw it away, and then the whole session is written to cache again at the higher rate. Decide those at the start.

Shorter replies, for the right reason. Ask for the answer without the essay. It helps, because the essay gets carried.

On a paid plan, /usage now shows which skills, subagents and MCP servers used your allowance. Worth a look before changing anything.

The tool that printed everything above is called baggage. It's free, it reads the transcripts already on your machine, and nothing leaves it. The name on npm belongs to someone else, so install it from the repo:

$ npm install -g github:manpreet171/baggage
$ baggage          # this project
$ baggage --all    # every project

What it doesn't tell you, so you don't over-read it:

Tokens, not dollars. The totals are the exact counts the API reported. Pricing them is your model and your plan.

Item sizes are estimates. The list ranks things with a rough four-characters-a-token ruler, the same ruler for everything. The totals above it are exact.

Only what's still on your machine. Claude Code keeps about 30 days of history by default.

One heavy user. These are my numbers. Yours will differ, which is the point of running it.

This is how we work. Measure what's really happening before changing anything, then change the thing the numbers point at. It's the same rule we use when an AI system we build for a business starts costing more than it should.

Sources: both terminal blocks are real baggage runs on my own machine on 5 Oct 2026, the second trimmed to whole lines for length. Prices and cache multipliers are from Anthropic's pricing page; the cost split weights my exact token counts by them. Cache and usage behaviour is from Claude Code's costs and prompt caching docs.

Read next: The best CLAUDE.md rules are hiding in your chat history · I researched loop engineering to build a product. I built nothing.

Originally published at singhlabs.dev.

── more in #ai-tools 4 stories · sorted by recency
── more on @claude code 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/where-your-claude-co…] indexed:0 read:7min 2026-10-05 · —