Stop Comparing Model Prices. Measure Cost per Completed Work. A developer argues that comparing AI model prices by token cost is misleading and proposes measuring "accepted work cost" — total task cost (model tokens, tool fees, retries, and human review) divided by the number of accepted work products. The approach holds the surrounding system fixed and varies only the model across five real tasks (coding, research, writing, data, and recurring operations), tracking tokens, tool calls, retries, wall time, human minutes, and acceptance status. The author cites Fireworks' vendor-reported Ember-1 figures showing token use dropping from 49.3K to 29.9K (39%) with task score essentially flat at 0.751 to 0.753, but notes this alone is insufficient for choosing a model in a tool-using system. Model pricing is easy to compare. Completed work is harder. Fireworks reports a customer comparison in which total token use moved from 49.3K to 29.9K, a 39% reduction, while the task score moved from 0.751 to 0.753. That is useful evidence for token efficiency. It is not enough evidence for choosing a model in a real tool-using system. The missing unit is an accepted work product. A model response can be generated without being usable. A coding agent can return a patch that fails the next test. A research agent can produce a fluent summary with weak source coverage. An operations agent can create a list that has no owner or deadline. The model bill captures tokens. The task also consumes retries, tool calls, context, wall time, and human correction. I use four states: For an evaluation, I keep the surrounding system fixed. The same person defines the task. The same prompt sets behavior. The same tools provide evidence. The same acceptance line judges both candidates. I change one variable: the model. Then I run five real tasks: coding, research, writing, data, and recurring operations. I record tokens, tool calls, retries, wall time, human minutes, accepted status, reusable-artifact status, and the reason for rejection. The calculation is deliberately plain: model cost = input tokens input price + output tokens output price task cost = model cost + tool fees + retry cost + human review cost accepted work cost = task cost / accepted work products The denominator must be visible. A ten-dollar run that produces five accepted outputs is not the same as a ten-dollar run that produces one draft and a long repair cycle. The acceptance line must match the work. For code, I might require a reproduced bug, a narrow change, passing existing tests, one regression test, and no unrelated files. For research, I might require five primary sources, one claim per source, dates on current claims, named uncertainty, and a decision note. For operations, I might require a dated input, a reason for each exception, an owner, a deadline, and a record that can be reopened tomorrow. This is also why I read the comparison through the Agent Stack™: Human, Behavior, Tools, Memory, Agents, Models. The model is one layer. Tools add outside work. Memory adds retrieval and repeated context. Agents add steps and retries. Human review defines acceptance. Comparing only the last layer hides the path that creates the bill. The current sources are bounded. Fireworks' Ember-1 figures are vendor-reported and the model is presented as a Research Preview. OpenAI's evaluation guidance recommends task-specific tests and human judgment. Neither source proves the result for my workload. They make the next experiment clearer. Before changing a model, run the work-unit trial. Measure the cost of something a person can accept, reopen, and use. Canonical article: https://echonerve.com/stop-comparing-model-prices-measure-cost-per-completed-work/ https://echonerve.com/stop-comparing-model-prices-measure-cost-per-completed-work/