{"slug": "my-subagents-were-eating-48-of-my-tokens-so-i-gave-them-hard-budget-caps-by-not", "title": "My Subagents Were Eating 48% of My Tokens. So I Gave Them Hard Budget Caps — Enforced by Hooks, Not Dashboards", "summary": "A developer built subagent-budget, an open-source Python tool that enforces per-agent-type token and dollar budgets for Claude Code subagents via PreToolUse hooks, blocking over-budget launches with exit code 2 instead of merely displaying spend. The tool was motivated by a widely circulated measurement showing subagents consumed 48.1% of tokens across 455 sessions, with a median subagent first request of 47,117 tokens. Budget rules are glob-matched against agent description and subagent type, persist across sessions in a JSONL ledger, and can import spend from the compatible resume-budget-guard tool.", "body_md": "A few weeks ago a measurement made the rounds: developer aidiveyt logged a\n\nmonth of Claude Code usage — 455 sessions, 2,631 subagent runs — and found\n\nsubagents had consumed **48.1% of all tokens**. The median subagent's *first request* alone was 47,117 tokens. Nobody approved that spend. It just\n\n`Task` call at a time.\nSimon Willison's reaction was the whole community's reaction: *\"hard budget caps, please, now.\"*\n\nI had the same problem and the same reaction. So I built\n\n`subagent-budget`: per-agent-type token **and** dollar budgets that are\n\n**enforced by Claude Code hooks** — an over-budget subagent is refused at\n\nlaunch, with the reason shown to the agent. A dashboard that can't stop a\n\nlaunch is a suggestion. This is a cap.\n\n```\npip install subagent-budget\n\nsubagent-budget init --default-tokens 1000000 --default-usd 50\nsubagent-budget set-budget --pattern \"Explore*\" --tokens 200000 --usd 10\nsubagent-budget sync        # backfill from ~/.claude/projects transcripts\nsubagent-budget report\n```\n\nThen the enforcement part — one block in `~/.claude/settings.json`:\n\n```\n{\n  \"hooks\": {\n    \"PreToolUse\": [\n      {\n        \"matcher\": \"Task\",\n        \"hooks\": [{ \"type\": \"command\", \"command\": \"subagent-budget hook --event pre\" }]\n      }\n    ]\n  }\n}\n```\n\nNow every subagent launch goes through a budget check first. Under budget:\n\nit launches, and the hook prints the remaining headroom. Over budget: exit\n\ncode 2 (Claude Code's blocking-hook signal), and Claude sees exactly why:\n\n```\nsubagent-budget blocked launch of 'Explore auth code' (Explore):\ntoken budget exceeded: used 210,441 / 200,000 tokens\n```\n\nThat last part is the whole point. Prompt-level instructions (\"please don't\n\nspawn too many subagents\") don't survive contact with an agent mid-task. A\n\nhook that returns a non-zero exit sits *below the prompt layer* — the model\n\ncan't argue with it, negotiate with it, or forget it.\n\nThere's already a tool in this space — `subagent-ledger` — and it's honest\n\nabout what it is: a *display*. It shows token counts, clears at session end,\n\nand can't stop anything. `subagent-budget` is built on the opposite premise:\n\n`~/.config/subagent-budget/ledger.jsonl` and survives session ends. A\nbudget that resets every session isn't a budget.\nBudget rules are globs matched against both the agent's description and its\n\nsubagent type, first match wins — so `Explore*` can have a tight shared pool\n\nwhile everything else falls back to the default. Either dimension can be\n\nleft unlimited.\n\nIf you already track spend across `claude --resume` with\n\n[resume-budget-guard](https://github.com/hahahahahahahahah6/resume-budget-guard),\n\nthe ledger formats are deliberately compatible (`ts`, `kind`, `cost_usd`\n\nrecords). One command folds resume spend into the same caps:\n\n```\nsubagent-budget import-rbg\n# imports ~/.resume-budget-guard/ledger.jsonl as agent \"rbg:<session-key>\"\n```\n\nImported spend counts toward the matching budget rules, so a resume can no\n\nlonger silently reset what a subagent cap was guarding. Idempotent — re-run\n\nit anytime.\n\n`sync` records one ledger entry per Task\ncall using a configurable estimate (default 47,117 — the measured median).\nUse `record --tokens` with exact figures from `--output-format json` when\nit matters.`model_rates` or pass `--cost` for exact numbers.`record`.\nStdlib only, zero dependencies, MIT. `pip install subagent-budget` —\n\n[GitHub](https://github.com/hahahahahahahahah6/subagent-budget) ·\n\n[PyPI](https://pypi.org/project/subagent-budget/).", "url": "https://wpnews.pro/news/my-subagents-were-eating-48-of-my-tokens-so-i-gave-them-hard-budget-caps-by-not", "canonical_source": "https://dev.to/haoli/my-subagents-were-eating-48-of-my-tokens-so-i-gave-them-hard-budget-caps-enforced-by-hooks-not-9e0", "published_at": "2026-10-11 09:29:02+00:00", "updated_at": "2026-10-11 09:51:41.658278+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "mlops"], "entities": ["Claude Code", "subagent-budget", "subagent-ledger", "resume-budget-guard", "Simon Willison", "aidiveyt", "GitHub", "PyPI"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/my-subagents-were-eating-48-of-my-tokens-so-i-gave-them-hard-budget-caps-by-not", "markdown": "https://wpnews.pro/news/my-subagents-were-eating-48-of-my-tokens-so-i-gave-them-hard-budget-caps-by-not.md", "text": "https://wpnews.pro/news/my-subagents-were-eating-48-of-my-tokens-so-i-gave-them-hard-budget-caps-by-not.txt", "jsonld": "https://wpnews.pro/news/my-subagents-were-eating-48-of-my-tokens-so-i-gave-them-hard-budget-caps-by-not.jsonld"}}