What 509 steps of our own AI Programming actually cost A developer experimenting with multi-model routing for AI coding agents reported that 509 coding steps cost $8.99 in actual model spend, versus a $30.91 pricing baseline calculated at Opus rates — roughly 71% lower. The developer cautioned that this was not a clean A/B test and that some individual runs cost more when cheaper models failed and required retries or escalation, arguing that "the cheapest model isn't necessarily the cheapest option" and that cost per passing step matters more than cost per token. The approach is being built into a tool called CUPID. AI Programming feels cheap when you look at a single request. One request might cost maybe a few cents. But once an agent starts working through hundreds of steps, those costs add up quickly. Especially when every step goes to a frontier model. We’ve been experimenting with a different approach. We are trying to break coding work into smaller steps and use different models depending on what each step needs. Across 509 coding steps, our actual model cost came to $8.99 For comparison, if we tae the same token usage and price it entirely at Opus rates, the cost would have been: $30.91. That's roughly 71% lower. That sounds great on paper, but there’s an important caveat. This was not a clean A/B test where we ran the exact same 509 steps once with our routing setup and once entirely on Opus. The $30.91 figure is a pricing baseline: we took the same token usage and recalculated what it would have cost at Opus rates. So we don’t treat “71% lower” as a universal result. In fact, some individual runs actually cost more. That usually happened when a cheaper model failed, had to retry, or the task had to be escalated to a stronger model anyway. A low token price doesn’t help much if you need multiple attempts to get a passing result. That led us to a more useful way of thinking about AI coding cost: The cheapest model isn’t necessarily the cheapest option. What matters is the cost of getting a step completed successfully. A more expensive model that gets the task right on the first try can sometimes cost less than a cheaper model that needs retries. On the other hand, paying frontier-model prices for every simple step is also wasteful. So the question we’ve been exploring is: What is the cheapest model that can reliably complete this specific step? That’s also why we’ve started looking beyond cost per token and toward cost per passing step. We’re still collecting more runs, especially the cases where routing performs worse. Those failures are probably more useful than the headline savings number, because they show us where cheaper models stop being worth it. I’d be curious how others are measuring this. Do you mostly track total API spend, or do you look at cost at the individual task level? We’re building this approach into CUPID https://cupidcode.ai https://cupidcode.ai . Happy to share more of the raw numbers in the comments.