cd /news/ai-tools/stop-comparing-model-prices-measure-… · home › topics › ai-tools › article
[ARTICLE · art-144299] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=· neutral

Stop Comparing Model Prices. Measure Cost per Completed Work.

A developer argues that comparing AI model prices by token cost is misleading and proposes measuring "accepted work cost" — total task cost (model tokens, tool fees, retries, and human review) divided by the number of accepted work products. The approach holds the surrounding system fixed and varies only the model across five real tasks (coding, research, writing, data, and recurring operations), tracking tokens, tool calls, retries, wall time, human minutes, and acceptance status. The author cites Fireworks' vendor-reported Ember-1 figures showing token use dropping from 49.3K to 29.9K (39%) with task score essentially flat at 0.751 to 0.753, but notes this alone is insufficient for choosing a model in a tool-using system.

by read2 min views1 publishedOct 3, 2026

Model pricing is easy to compare. Completed work is harder.

Fireworks reports a customer comparison in which total token use moved from 49.3K to 29.9K, a 39% reduction, while the task score moved from 0.751 to 0.753. That is useful evidence for token efficiency. It is not enough evidence for choosing a model in a real tool-using system.

The missing unit is an accepted work product.

A model response can be generated without being usable. A coding agent can return a patch that fails the next test. A research agent can produce a fluent summary with weak source coverage. An operations agent can create a list that has no owner or deadline. The model bill captures tokens. The task also consumes retries, tool calls, context, wall time, and human correction.

I use four states:

For an evaluation, I keep the surrounding system fixed. The same person defines the task. The same prompt sets behavior. The same tools provide evidence. The same acceptance line judges both candidates. I change one variable: the model.

Then I run five real tasks: coding, research, writing, data, and recurring operations. I record tokens, tool calls, retries, wall time, human minutes, accepted status, reusable-artifact status, and the reason for rejection.

The calculation is deliberately plain:

model_cost = input_tokens * input_price + output_tokens * output_price
task_cost = model_cost + tool_fees + retry_cost + human_review_cost
accepted_work_cost = task_cost / accepted_work_products

The denominator must be visible. A ten-dollar run that produces five accepted outputs is not the same as a ten-dollar run that produces one draft and a long repair cycle.

The acceptance line must match the work. For code, I might require a reproduced bug, a narrow change, passing existing tests, one regression test, and no unrelated files. For research, I might require five primary sources, one claim per source, dates on current claims, named uncertainty, and a decision note. For operations, I might require a dated input, a reason for each exception, an owner, a deadline, and a record that can be reopened tomorrow.

This is also why I read the comparison through the Agent Stack™: Human, Behavior, Tools, Memory, Agents, Models. The model is one layer. Tools add outside work. Memory adds retrieval and repeated context. Agents add steps and retries. Human review defines acceptance. Comparing only the last layer hides the path that creates the bill.

The current sources are bounded. Fireworks' Ember-1 figures are vendor-reported and the model is presented as a Research Preview. OpenAI's evaluation guidance recommends task-specific tests and human judgment. Neither source proves the result for my workload. They make the next experiment clearer.

Before changing a model, run the work-unit trial. Measure the cost of something a person can accept, reopen, and use.

Canonical article: https://echonerve.com/stop-comparing-model-prices-measure-cost-per-completed-work/

── more in #ai-tools 4 stories · sorted by recency
── more on @fireworks 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/stop-comparing-model…] indexed:0 read:2min 2026-10-03 · —