{"slug": "fable-decides-opus-and-sonnet-do-the-work-how-i-route-claude-code-subagents", "title": "Fable Decides, Opus and Sonnet Do the Work: How I Route Claude Code Subagents", "summary": "A developer's Claude Code subagent harness routes planning and final review to Fable while delegating execution to Opus, Sonnet and Haiku, after a June workflow script with a single model-less agent() call spawned 110 copies of Fable and consumed an entire five-hour usage window. A September 17 audit counted 334 million tokens of cache reads across 30 sessions and 1,943 turns, with a Fable session starting at 77k to 83k tokens before any input, and a September 2 measurement put a lean agent at about 17k tokens to arrive versus about 60k for a general-purpose agent. The setup pins each of four agent types — opus-owner, sonnet-implementer, haiku-scout and advisor — to a specific model with a restricted tool list, and requires every worker to close with a footer listing what changed, the verifying command, paths checked and what was not checked.", "body_md": "In June one workflow script had a single `agent()` call with no model on it. Subagents inherit the parent's model, the parent was Fable, and that one line spawned 110 copies of Fable. The run ate my entire five-hour usage window by itself.\n\nI see a lot of people blaming Fable for their usage limits right now. Some of that is fair, it's the most expensive model I run. But most of what I see is the same mistake I made in June. Fable doing work Opus or Sonnet should be doing, in a context Fable then has to re-read on every turn.\n\nSo my Claude Code harness treats Fable like the person running the job, not the person doing it. Fable plans, decides, reviews and writes the final answer. Opus owns the big chunks. Sonnet writes code once the decision is made. Haiku looks things up. Hooks enforce most of that, and the part hooks can't enforce is a short briefing Fable gets once per session.\n\nHere's how it's wired, what each piece costs, and the two times I got the balance wrong in opposite directions.\n\n## Why the top seat is the expensive one\n\nThe bill in an agent session isn't what the model writes. It's the context it re-reads every turn. My September 17 audit counted 334 million tokens of cache reads across 30 sessions and 1,943 turns. A Fable session in my home directory starts at 77k to 83k tokens before I've typed anything.\n\nAnything a tool returns lands in that context and stays there. A test run, a 40-file grep, a build log. Every later turn re-reads it, and in a Fable main loop it gets re-read at Fable prices.\n\nSo the goal isn't really \"use Fable less.\" It's keep bulky output out of Fable's context. Fable should see decisions and conclusions. The raw material goes to a cheaper context, gets worked there, and comes back as a short report. When the worker finishes, its context gets thrown away.\n\nOne thing worth saying up front. My rules put Opus in the main loop for routine work, not Fable. Fable gets the seat for architecture calls, security review, work that spans several projects, and a root cause after two fixes have already failed. It also shows up at the end of big workflows as the judge, which gets its own section below. Everything here is about what happens once Fable is in the seat.\n\n## The org chart\n\nThere are four agent types, each one a markdown file with a pinned model and a short tool list.\n\n- **opus-owner** runs on Opus. It owns a large or risky task end to end: a multi-file change, a repo sweep, a root cause hunt. It's the only worker with the Agent tool, so it's the only one that can hand pieces further down.\n- **sonnet-implementer** runs on Sonnet. The decision is already made, it writes the code inside the scope it's given. In a fix loop it's not allowed to edit test files. A test that looks wrong comes back up as a decision instead of getting bent until it passes.\n- **haiku-scout** runs on Haiku. Read-only lookups: where is this defined, what calls it, how many are there. If it can't find something it has to say NOT FOUND and list where it looked.\n- **advisor** is read-only, answers one question in under 400 words, and spawns nothing. More on that one below.\n\nNone of them get skills, MCP servers or the artifact tools. That's on purpose. On September 2 a lean agent cost about 17k tokens to arrive and a general-purpose one about 60k, before either one did anything. I wrote [a whole post on measuring that](https://www.practicalsystems.io/blog/measure-your-subagent-token-cost), so the short version: a scout running a grep doesn't need the catalog of every design skill I've installed.\n\nThe other thing they share is how they end. Every worker closes with the same footer: what changed, the exact command that verified it, the paths it checked, and what it didn't check. A claim with no evidence line counts as unverified. That footer is what Fable reads. Not the transcript.\n\nHere's what they actually cost. The harness has measured every spawn since September 16, and across 86 of them the median runs looked like this:\n\n| Worker | Median tool calls | Median tokens per run | \n|---|---|---|\n| haiku-scout | 1 | 17,533 | \n| sonnet-implementer | 53 | 493,723 | \n| opus-owner | 110 | 2,115,498 | \n\nThose are effective tokens, with cache reads counted at a tenth. The median Opus owner made 110 tool calls of reading, editing and testing inside its own context. None of those 110 results ever sat in Fable's context. Fable got the footer.\n\nThe same ledger records who started each one. Fable started 56 of the 86: 33 Opus owners, 5 Sonnet implementers, a couple of one-offs, and 16 general-purpose agents, which is the heavy type and more of them than I'd like. Opus parents started 15 Sonnet implementers and 4 Haiku scouts. Sonnet parents started 6 Haiku scouts.\n\n## Work goes down, questions go up\n\nWho can spawn whom is a graph, and a hook checks it on every spawn.\n\nFable can spawn Opus, Sonnet or Haiku. Opus can spawn Sonnet or Haiku. Sonnet can spawn Haiku. Haiku spawns nobody. Delegation only goes down. Peers aren't edges either, because a Sonnet worker handing its job to another Sonnet worker buys nothing and just hides the work one level deeper. A fork counts as a peer for the same reason, since it inherits the caller's model.\n\nThe one thing allowed to go up is a question. Any agent above Haiku can spawn the advisor, and the advisor always runs one rung above whoever called it. Sonnet asks Opus. Opus asks Fable. Fable asks Fable. The hook ignores whatever model the caller asked for and injects the right one. A peer advisor has the same blind spots you do, and a Fable advisor for a Sonnet worker burns Fable where Opus would have been fine.\n\nThe advisor never takes the task over. It answers and hands it straight back. Consultations aren't capped either, because a blocked question turns into a guess, and a guess in the middle of a fix pass costs more than the advice would have.\n\nBetween September 2 and 16 that hook logged 24 advisor consultations and blocked 25 spawns: 22 that went upward or sideways, and 3 where a Haiku worker tried to spawn something.\n\nTwo more rules sit on top. Depth two is the ceiling, so Fable to Opus to Haiku and never deeper. And fan out from the highest level that can already write the brief. If Fable already knows the five files, it spawns five Haiku scouts, not one Opus owner that goes and rediscovers them.\n\nClaude Code has its own built-in advisor now, but it's one global setting with no per-agent override, so it can't escalate relative to the caller. I [filed an issue](https://github.com/anthropics/claude-code/issues/91715) asking for that. Until it lands my Opus workers are told to skip the built-in one, since it runs on Opus too and would just be a peer.\n\n## What Fable writes when it hands work down\n\nEvery spawn has to name its model and declare its size. The declaration is one line in the prompt:\n\n```\n# EST: 14 calls, 3 files\n```\n\nA hook reads that line and prices the spawn two ways. Delegated, it's the arrival cost plus a couple thousand tokens per call. Inline, it's every result landing in the main context and getting re-read for the rest of the session. If the spawn doesn't pay for itself it gets denied, and the deny message shows the arithmetic.\n\nTwo real entries from September 16. An Opus owner declared a 25-call read-only audit: about 67,000 tokens to spawn against about 150,000 inline, so it went through. A general-purpose agent on Opus for one tool call: 62,000 to spawn against 6,000 inline, denied.\n\nThe rule of thumb that falls out of it is under about ten tool calls or eighty edited lines, do it inline, no matter who's in the main loop. Delegation earns its arrival cost when the work is big, when the pieces can run at the same time, or when the output is long and you only need the conclusion.\n\nA full dispatch from the Fable loop looks roughly like this. It's an example, not a log line:\n\n```\nAgent({\n  subagent_type: \"sonnet-implementer\",\n  model: \"sonnet\",\n  description: \"Retry failed webhook sends\",\n  prompt: `# EST: 14 calls, 3 files\nScope: src/webhooks/send.ts, src/webhooks/retry.ts, test/webhooks.test.ts\nDone when: a 503 retries three times with backoff and the existing tests pass\nVerify: npm test -- webhooks\nReport: files changed, the verify command and its last line, what you did not check`\n})\n```\n\nFable decides the scope, the finish line and the check. Sonnet does the typing. Fable reads four lines back.\n\nOne detail in that hook matters more than the math. The identical dispatch gets denied at most once. The second attempt goes through and the reason gets logged. That came straight out of the next section.\n\n## The guard I had to take out\n\nBefore any of this, I tried to force it. I wrote a hook that blocked Fable from doing hands-on work, with a per-prompt edit budget and a denial on any shell command that wrote code. If Fable wanted to edit, it had to delegate.\n\nIts log hit 1,046 events. 714 of them were overrides. It denied 266 shell commands, including `npm test`, a commit message written with a heredoc, and a read-only grep. And it denied 47 edits that showed up in runs of five to seven on the same file, because the model treats a deny like a flaky error and just retries the next queued edit. Once the cap fired at edit 21 of a change that belonged together and left a file half-edited, which is worse than finishing it or never starting.\n\nThe other side showed up the same week. On September 3, the day I launched declick, three delegated fix passes cost 3.7 million tokens and two hours on defects that a 40-minute pass by hand closed. The budget was fighting the hand pass.\n\nSo on September 6 I took the block out. The hook does two things now. It briefs a Fable session once, at the start, with the measured numbers: what a spawn costs, where the break-even sits, the one-line edit that cost 77k when it got delegated. And it logs the hand work Fable does anyway, so I can look at it later. It denies nothing.\n\nWhat I took from it is that a hard block in the middle of a task gets probed, not obeyed. Block the things that destroy state. For routing, brief up front and fix the call on the way through. That's where the next piece came from.\n\n## Routing moved into Claude Code itself\n\nClassic hooks are processes. On September 19 I measured a single Bash call firing 15 PreToolUse hooks and 4 PostToolUse hooks, each one its own process spawn, and on this machine a spawn can take over a second. Every subagent brings that whole chain along with it.\n\nOn September 16 routing moved into Claude Code's function hooks, which run inside the Claude Code process instead of spawning anything. I call that layer Mods, and [the post about it](https://www.practicalsystems.io/blog/claude-code-function-hooks-mods-layer) is mostly about how it sat there enabled and doing nothing for a while. It works now.\n\nThe bigger change isn't speed though. The Mod rewrites instead of denying. If a spawn asks for the wrong model, the Mod moves it onto the graph and tells the parent what it changed and why, right on the Agent tool result. 8 of the 86 measured spawns got rewritten. Two haiku-scouts that were asked for on Opus ran on Haiku. An advisor asked for on Haiku ran on Opus, one rung above its caller. The only things it still denies are a Haiku worker trying to spawn, and a dispatch with no size declared on its first try.\n\nIt also measures every spawn and learns a median per agent type. The first haiku-scout it saw was reserved at 17,000 tokens and came in at 17,006. That learned number feeds the break-even math now, instead of a constant I measured once by hand.\n\nThe old hooks are still installed. If the Mod stops checking in for a session, they take over.\n\nSame idea for the rules themselves. The full delegation contract is about 1,350 tokens, and Fable doesn't carry it every turn. A context hook loads it only when a prompt touches things like subagents or spawning, and says why it loaded. The session I used to draft this post loaded it because my request said \"subagents.\"\n\n## Workflows: Fable only at the end\n\nThe biggest jobs run as Workflow scripts. A script fans out to a lot of agents at once, and then something has to read everything they found and make the call. That last step is the one place Fable gets spawned as a subagent.\n\nThe rules for that are strict because of June. Every `agent()` in a script names its own model inline, and one bare call blocks the whole script. Fable is allowed at most three times per script, and only as a top-level `await` after the fan-out. Never inside a `parallel`, a `map` or a loop, where one line can turn into a hundred agents.\n\nThen September 19 showed me the part I was missing. A workflow auditing my own harness ran 12 Opus readers, then one Sonnet skeptic per duplicate the readers claimed. The readers came back with 212 claims that overlapped heavily, with the same mirrored file reported by five different readers. The run planned around 225 agents and got to 186 before I noticed my computer had slowed to a crawl. Of the 22 verifications that finished, none changed the decision. The right size was about 25.\n\nNow every script declares its size as arithmetic before it runs: producers, plus items times verifiers, plus synthesizers. Any term that depends on an earlier stage gets capped with a slice, and the drop gets logged. Duplicates get merged before any per-item stage. The ceiling is 30 agents, anything over 40 needs a written reason, and a fan-out with no declared size gets denied.\n\n``` js\n// MAX_AGENTS: 12 readers + 10 verifiers + 1 judge = 23\nconst found = await parallel(AREAS.map(a => () =>\n  agent(readBrief(a), { model: 'opus', agentType: 'opus-owner' })))\nconst claims = dedupe(found.flat())\nconst checked = claims.slice(0, 10)\nlog(`verifying ${checked.length} of ${claims.length} claims`)\nconst verdicts = await parallel(checked.map(c => () =>\n  agent(verifyBrief(c), { model: 'sonnet' })))\nconst report = await agent(judgeBrief(verdicts), { model: 'fable' })\n```\n\nConcurrency doesn't save you here either. The 16-slot limit only queues the rest. Every one of them still runs, and every one of them spawns the full hook chain on my desktop.\n\n## The bug this post found\n\nWhile I was drafting this post, the Mod blocked the two Sonnet workers I sent to draw these figures. The deny said a Sonnet spawn costs 487,680 tokens before its first tool call. The measured arrival for a lean worker is about 17,000. So that number was off by roughly 28x, and it turned out to be the median of whole Sonnet runs, not the arrival.\n\nSo I went and looked. The Mod is supposed to record what a worker spends before its first tool call. Of the 34 runs logged since that field was added, it captured that number 0 times. The tool call gets counted before the step that paid for it shows up, so the arrival never lands, and the fallback stores the whole run instead. Every Opus owner was getting priced at about 2.2 million tokens just to show up.\n\nThe only reason it didn't block everything is the rule from the guard I took out: the same dispatch gets denied once, then goes through. And the Fable loop had been routing around it quietly. 31 of the 56 spawns Fable made that week went through on an override.\n\nThe fix was a few lines. Count the first step as the arrival no matter when its tool call gets counted, never store a whole run as overhead, and throw out the old samples since every one of them was a total. I wrote the failing test first, watched it fail, then fixed it. A fresh session now prices a Haiku scout at 17,000 again instead of 25,580.\n\nThat's the same lesson as the rest of this post, pointed at me. A learned number that was never checked against a known answer is a guess with a log file.\n\n## What to take\n\nIf you're running Fable in Claude Code and hitting limits, this is what I'd copy, in order.\n\n**Name the model on every spawn.** That one rule would have saved the June window by itself, and it's a few lines of hook.\n\n**Keep bulky output out of Fable's context.** Fable decides, reviews and synthesizes. Work that throws off a lot of tool output goes to a cheaper worker that reports back with evidence.\n\n**Delegate by size, not by principle.** Under about ten tool calls, stay inline. Measure your own arrival cost first, because mine won't match yours.\n\n**Brief and rewrite, don't wall.** A mid-task block gets probed. Put the economics in front of the model once and fix the routing on the way through.\n\n**Size the fan-out before you launch it.** Write the agent count down as arithmetic. If you can't, cap it.\n\nThe whole harness is public at [claude-harness](https://github.com/ucsandman/claude-harness), a mirror of [Agnostic AI](https://github.com/ucsandman/Agnostic-AI). The agent definitions, the routing hooks, the Mods layer and the workflow guard from this post are all in there. If you only take one piece, take the capability graph.", "url": "https://wpnews.pro/news/fable-decides-opus-and-sonnet-do-the-work-how-i-route-claude-code-subagents", "canonical_source": "https://www.practicalsystems.io/blog/fable-opus-sonnet-claude-code-subagent-routing", "published_at": "2026-09-29 04:49:15.859958+00:00", "updated_at": "2026-09-29 04:49:17.763565+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "large-language-models", "developer-tools", "mlops"], "entities": ["Claude Code", "Fable", "Opus", "Sonnet", "Haiku", "opus-owner", "sonnet-implementer", "haiku-scout"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/fable-decides-opus-and-sonnet-do-the-work-how-i-route-claude-code-subagents", "markdown": "https://wpnews.pro/news/fable-decides-opus-and-sonnet-do-the-work-how-i-route-claude-code-subagents.md", "text": "https://wpnews.pro/news/fable-decides-opus-and-sonnet-do-the-work-how-i-route-claude-code-subagents.txt", "jsonld": "https://wpnews.pro/news/fable-decides-opus-and-sonnet-do-the-work-how-i-route-claude-code-subagents.jsonld"}}