In my Claude Code sessions, 69% of the output came from the top models (Opus, Fable), including reading a changelog, applying an edit I'd already decided on, and running tests.
"Hard task → strong model" didn't fix it. Difficulty is the wrong signal: once a refactor is decided, applying it needs no judgment, however complex the code. What worked was routing by the kind of decision a step needs:
I built this as advisor-mode in a plugin called maddog. Your main session splits the goal and sends each piece to a subagent (a separate Claude instance with its own context) on the right model. Only a short result comes back, and the main session checks it before accepting. Merges and pushes need your go-ahead. That's a rule the agents follow, backed by a hook in subagents, not a hard lock.
My numbers, from my own 49 advisor-mode sessions against 30 sessions without it over the same weeks (output tokens, not cost or quality):
In my use, small changes cost more to dispatch than to just do, and Haiku fails on tasks that only look mechanical.
The plugin also has section-by-section, which reviews a skill or agent file with you one section at a time.
Video (~100 s): https://youtu.be/b_mFbVY5jzk
Try it: in Claude Code, run /plugin and search "maddog" on the Discover tab, then start with /maddog:advisor-mode <your task>. Run the main session on Sonnet or above. MIT, repo: https://github.com/Harish-here/maddog
Do you route work across models, or run everything on one?