cd /news/ai-agents/delegating-code-to-claude-did-not-au… · home › topics › ai-agents › article
[ARTICLE · art-141466] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Delegating code to Claude did not automatically save Codex tokens

A developer built TaskRoute, a local Codex plugin that delegates Python function implementation to Claude Code while Codex verifies the returned work. In a single sequential comparison on a synthetic Python task, the delegated route cut Astra's uncached input tokens by 59% and output tokens by 83% with elapsed time nearly unchanged, though earlier runs showed delegation saving as little as 3% and taking 49% longer. The developer cautions the results are one non-randomized pair and do not establish equivalent savings in subscription allowance.

by read3 min views6 publishedSep 29, 2026

I use Astra in Codex for work beyond software development. Opus writes code well, so I wanted to hand implementation to Claude Code and keep more of Astra's allowance for everything else.

The obvious workflow was to send the task, wait for the code, and check the result. I started building TaskRoute to handle that exchange without making me copy messages between two models.

I compared that route with having Astra implement the same task directly.

In one early comparison, delegation reduced Astra's uncached input tokens by only 3% and took 49% longer. Both versions passed the same 45 independent checks.

An even smaller task had been worse: 29% more uncached input, despite producing less output.

The logs showed the work Astra was still doing around the handoff. Astra read large snapshots, repeated verification, and assembled the result from several files. In the 45-check comparison, the delegated route needed 10 model steps versus 7 for direct implementation.

Removing code generation had removed only part of Astra's work.

I collected the changes, test results, and independent reviewer findings into one compact acceptance packet. Astra still had to inspect the actual result. The packet made that inspection less scattered.

Waiting needed attention too. One compact-handoff run still woke Astra 13 times while Claude worked. A later run reduced that to two waits by letting the process wait longer between returns to the model.

That change reduced repeated context processing, but it wasn't a controlled measurement of waiting alone. The reviewer output and model behavior also varied.

I then compared both routes on a new synthetic Python task: planning batches of dependent tasks, including priorities, completed tasks, missing dependencies, and cycles.

Both routes started with the same stub and contract, used fresh Astra sessions, and faced 20 frozen acceptance tests. Each implementation also added eight tests.

Astra execution workload Change with TaskRoute
Uncached input tokens 59% lower
Output tokens 83% lower
Input including cache 44% lower
Elapsed time Almost unchanged

Both implementations passed all 20 common tests and their eight additional tests. No live retries were needed in this pair.

These percentages describe Astra's execution sessions. Claude's usage is separate. Shared experiment preparation is also separate. This was one sequential pair on a new task, with visible tests, not a broad or randomized benchmark. It does not establish equivalent savings in subscription allowance.

In an earlier run, a missing temporary directory stopped the process before review, although Claude had already produced code.

The successful repeat alone showed a 34% reduction in uncached input. Including the failed attempt reduced that saving to 7%, before counting the parent session's diagnosis and repair work.

I added a local preflight to check the required files and directories before calling the models.

Another comparison could not establish equal quality. Both implementations passed their own tests, but the shared checks exposed different interpretations of which error to return when several rules failed.

The task description had left that precedence ambiguous. I clarified the contract instead of treating the result as a model ranking.

TaskRoute is a small local Codex plugin for bounded Python function changes on macOS, using an existing Claude Code setup. Claude implements and brings in a separate reviewer; Codex checks the returned work.

The repository includes the measured comparison and its limits, plus installation instructions for Codex.

I still need to measure how this affects an ordinary working week. For now, the useful change is concrete: Astra receives a result it can inspect without spending so much effort managing the exchange.

If TaskRoute is useful to you, GitHub stars are welcome.

── more in #ai-agents 4 stories · sorted by recency
── more on @taskroute 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/delegating-code-to-c…] indexed:0 read:3min 2026-09-29 · —