Everyone Measures AI Usage. 70% Can't Measure What It Returned. Anthropic's internal survey of 132 engineers found that Claude Code boosted merged pull requests per day by 67 percent and daily usage from 28 to 59 percent, with self-reported productivity gains of 20 to 50 percent, yet the company's delivery dashboard showed no movement in delivery metrics. The finding highlights a broader measurement gap: McKinsey found only 30 percent of leaders could say where AI-freed time went, and a Gartner 2025 survey found just 22 percent of leaders reported significant value from their AI tools. The analysis argues that adoption and token spend are costs rather than returns, and that business cases typically inflate value by counting soft hours as hard dollars and calendar time as labor time. Anthropic surveyed 132 of its own engineers about Claude Code. Merged pull requests per day rose 67 percent. Daily use of the tool climbed from 28 to 59 percent. Self-reported productivity gains ran between 20 and 50 percent. But then someone checked the organization's delivery dashboard and saw that the delivery metrics had not moved. That gap is the whole subject of this piece. A tool can be used constantly, rated highly by the people using it, and leave no trace on the numbers a business actually runs on. The measurement problem underneath it is bigger than one company's coding assistant. McKinsey found that 30 percent of leaders could say where the time AI freed up actually went. The other 70 percent could not. Seven out of ten organizations have people spending less time on tasks and no idea whether that turned into anything. Ask the question from the other side and the answer is just as thin: in Gartner's 2025 survey, 22 percent of leaders said their AI tools had returned significant value, a share that lands where McKinsey, Deloitte and ServiceNow each arrived measuring it their own way. Usage and tokens are costs, not returns Two numbers get reported as if they answered the ROI question, and neither does. Adoption tells you whether anyone is using the thing. It is a leading indicator and a useful one, but a tool with high adoption and no measured outcome has produced activity, not value. Token spend tells you what the tool costs to run. It belongs in the calculation, on the cost side, and watching it closely tells you nothing about whether the work it produced was worth having. Both are easy to pull from a dashboard, which is exactly why they get reported. The number that matters sits one step further out and takes real work to produce: whether the company made or saved a defensible dollar. Everything below is how you get to that number without lying to yourself on the way. So when a business case lands on a desk claiming a figure in saved dollars, the honest question is not whether AI helped. It is whether the number is real. In most cases I have looked at, it is inflated, and usually in two separate places at once. Inflation one: soft hours counted as hard dollars Here is the calculation almost every AI business case runs. The tool saved each person two hours a week. Multiply the hours by the hourly cost of those people, add it up across the team, and report the total as money saved. The hours are usually real. The dollars usually are not, because the budget did not change. Nobody was let go, no contractor was dropped, no line item fell. What happened is that a group of salaried people have slightly lighter weeks, and the company pays them exactly what it paid before. Freed hours become real money in three situations: the time is redeployed onto work that generates value, or it lets you avoid a hire you were about to make, or it lets the same headcount produce more of something you sell. If none of those is true, the saving is soft, and soft savings do not survive contact with a CFO who can see the budget did not drop. This is where the 70 percent from the opening returns. If seven in ten leaders cannot say where the freed time went, then most of the "hours saved times hourly rate" figures in circulation are soft hours nobody traced to an outcome, dressed up as hard dollars. The fix is not complicated. Label every dollar of claimed value as hard or soft, and report the two separately. The number gets smaller, and much harder to dispute. Inflation two: calendar time counted as labor time The second inflation is subtler, and I see it most in engineering cases. A feature "took three weeks" before and "takes one week" now, so the case dollarizes two weeks of saved time at an engineer's rate. The problem is that three weeks was never three weeks of work. Some of it was a ticket sitting in a queue, some was waiting on a review, some was a dependency that had not shipped. Cycle time, the calendar span from request to delivery, is not the same as labor time, the hours a person actually spent. Converting the calendar span to dollars at an hourly rate invents labor that no one performed. Speed is still worth reporting, but as a rate: this class of work now moves through 40 percent faster. It becomes money only when moving faster captures something real, most often revenue that arrives earlier because the thing shipped sooner. When it does not, a faster cycle is still a genuine improvement worth reporting as speed, but it is not a number you convert into dollars. There is a related trap in trusting self-reported speed at all. The one controlled study I know of that timed the same developers with and without AI, METR, found a measured slowdown of 19 percent against a self-reported gain of 20 percent. The developers were sure they were faster. The clock disagreed. Whatever you build your ROI on, it should not be a survey asking people how much time they think they saved. What actually produces dollars Strip out the inflations and there are only two mechanisms by which an AI tool produces money, and each converts to dollars through a different bridge. The first is acceleration: someone does a task they already did, in less time. The bridge is hours saved multiplied by the hourly cost of that person. If a task dropped from four hours to two and a half and happens eighty times a month, that is 120 hours a month, and at their hourly cost you have a real figure. Then you subtract rework, because generated output a person has to redo never saved the time it appeared to. The second is avoided work: a task stops happening at all. A support ticket the knowledge base resolves is a ticket a human never touches. The bridge here is not an hourly rate, it is the full cost of one whole interaction: the total monthly cost of the function divided by the number of interactions it handles. If support costs 20,000 a month and handles 2,500 tickets, each avoided ticket is worth 8 dollars, and that 8 already carries the tooling and overhead, not just one agent's wage. Using the acceleration bridge here, an hourly rate, would undercount it. Both of those are illustrative arithmetic, not figures from any client. The point is the shape: pick the mechanism, pick the matching bridge, and do not mix them. The denominator nobody writes down A return needs a cost to divide by, and this is where the tokens finally belong. The total cost of an AI tool per month is the build cost amortized over the months it will run, plus token spend, plus infrastructure, plus maintenance, plus any human review of its output. A build that took 150 hours and will run two years is not a 9,000-dollar hit this month; it is a few hundred a month spread across its life. Token cost is simpler than it looks: measure the average number of tokens one output consumes, multiply by the price per token, and you have a stable cost per output to hold against the value that output produces. A project can carry many metrics, but it has a single ROI. Each metric adds its own slice of value to the same numerator, and every slice divides by that same total cost. The one discipline this requires is avoiding double-counting: if a token cost already sits inside a per-unit figure, it does not also go in the denominator, and if two metrics describe the same saved dollar from two angles, you keep one of them. The math is deterministic, which is why it gets skipped I argued in an earlier piece that AI belongs at the edges of financial analysis and never in the arithmetic itself. Measuring your own AI's return is the same shape seen from the other side. The calculation is deterministic: a subtraction, a multiplication, an hourly cost, a total. No model is required, and none should be trusted with it. The reason it gets skipped is not difficulty. It is that the honest number is almost always smaller than the inflated one, and smaller numbers are harder to carry into a budget meeting. A defensible small number survives scrutiny and an impressive large one does not, and the second time a leader is caught reporting soft hours as hard dollars, the whole program's credibility pays for it. But isn't agentic AI supposed to be exempt from ROI? There is a serious version of the opposite argument, and Gartner makes it: early agentic AI is experimental, and organizations that demand a proven business case before they will touch it risk being outpaced by the ones that treat it as something to iterate on. That is right, as far as it goes. You do not gate a two-week experiment behind a formal ROI model, and pretending you can forecast the return on something genuinely new is a guess dressed as a forecast. But two different claims get folded together under that banner. "Do not require a business case before you experiment" is defensible. "Do not measure what it returned" is not, and the first is routinely used to justify the second. An experiment you never measure is not an experiment, it is a purchase with no follow-up. The reason for not gating early work behind ROI is to buy yourself room to find the value, which only means something if you then check whether you found it. Exempting agentic AI from a business case at the start is reasonable. Exempting it from measurement forever is how the share of companies that see real value stays stuck at one in five. Where to start Capture the baseline before you deploy anything, because once the old way is gone you cannot reconstruct how long it used to take. Measure one real unit of work end to end, from request to delivered, including the waiting, and see whether that number moved. Stop reporting seats deployed, tasks completed, prompts submitted and self-reported speed, and start reporting the one figure that ties to a customer or a budget. Label every claimed dollar hard or soft, and report the net dollars alongside the percentage, because the percentage moves with whatever you put in the denominator while the net figure does not. Doing that across every process you run, and writing down what each would need in order to actually change a budget line, is a workflow audit. It takes longer than pulling a usage chart. It is also the only version of the exercise that produces a number you can defend. Questions this raises How do I put a dollar value on time saved? Multiply the hours saved by the hourly cost of the person who saved them. Count it as a real saving only if that time is redeployed, avoids a hire, or produces more sellable output. Otherwise it is a soft number and should be labelled as one. Should token spend count as part of ROI? Yes, as a cost, in the denominator, and never as a benefit. Track cost per unit of useful output if you want a figure to hold against value, and make sure you are not counting the same tokens in two places. Can one project have several ROIs? No. A project has one ROI. Several metrics can each add a slice of value to the same numerator, divided by the same total cost. If you end up with two ROIs for one tool, you have either double-counted or mixed two projects together. https://unlockedconsulting.ai/blog/everyone-measures-ai-usage-70-can-t-measure-what-it-returned https://unlockedconsulting.ai/blog/everyone-measures-ai-usage-70-can-t-measure-what-it-returned