{"slug": "ais-measurement-crisis-is-over-the-translation-crisis-is-next", "title": "AI’s measurement crisis is over. The translation crisis is next", "summary": "MIT's GenAI Divide report found that 95% of enterprise generative AI pilots delivered no measurable P&L impact despite $30-40 billion in spending, but researchers at UC Berkeley argue the figure reflects poor measurement rather than failed projects. In the first half of 2026, enterprise AI investment pivoted to employee-facing use cases with existing KPIs, as Foundry's 2026 AI Priorities study shows 55% of IT decision-makers cite improving employee productivity as the top objective driving AI investment.", "body_md": "Last fall, you couldn’t open a business publication without tripping over some version of the same headline: where is the ROI for AI? The anchor for most of that coverage was [MIT’s “GenAI Divide” report](https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/), which found that despite $30 to 40 billion in enterprise generative AI spending, 95% of pilots delivered no measurable P&L impact. The bubble takes wrote themselves. Boards asked uncomfortable questions. More than a few AI budgets went into the freezer for the winter.\n\nHere’s the detail that got lost in the panic: the study defined success as measurable KPI impact within six months of the pilot. Read that again. A project that transformed how a team worked but was never instrumented to prove it counted as a failure. [Researchers at UC Berkeley pushed back](https://exec-ed.berkeley.edu/2025/09/beyond-roi-are-we-using-the-wrong-metric-in-measuring-ai-success/) on exactly this point, arguing that the 95% figure may represent 95% of organizations measuring the wrong things at the wrong time rather than 95% of projects failing to create value.\n\nIn other words, the AI ROI crisis of 2025 was never really about the AI. It was about measurable verification. Most enterprise AI projects didn’t fail. They were simply built in a way that made success unprovable. If you’re a CIO defending a budget line, that distinction is cold comfort, because “we can’t tell if it worked” and “it didn’t work” produce the same conversation with your CFO. But the diagnosis matters, because the treatment is completely different. You don’t fix an unprovable project with a better model. You fix it by picking a better problem.\n\nI’ve [argued before](https://thenewstack.io/theres-no-sku-for-ai-a-3-box-framework-to-avoid-ai-failures/) that AI initiatives should start with problems that already have good data and trusted metrics, and over the first half of 2026, the market arrived at that conclusion on its own.\n\nWatch where enterprise AI money actually went in the first half of this year and you’ll see a pattern that never made headlines: a hard pivot toward employee-facing use cases. Agents assisting support reps, sales teams, claims processors, IT help desks. The conventional read is that these are the safe choices, the training-wheels projects companies run while they work up the nerve for customer-facing AI.\n\nThat read is wrong. The pivot to employee-facing AI isn’t about safety. It’s about scoreboards.\n\nThink about what an employee-facing workflow comes with that a greenfield AI initiative doesn’t. You already measure it. Average handle time, first-call resolution, cases closed per week, quota attainment. Those KPIs have years of baseline data behind them. More importantly, they’re politically real. In many organizations, people are bonused on those numbers. Nobody in the room disputes the methodology of a metric that’s been sitting on a comp plan for five years. When you drop an agent into that workflow and the KPIs move in the right direction across the entire employee population, ROI stops being a philosophy seminar and becomes back-of-the-envelope arithmetic. Headcount, fully loaded cost, percentage improvement, multiply.\n\nThe survey data backs up what I’ve been seeing in the field. [Foundry’s 2026 AI Priorities study](https://foundryco.com/research/research-ai-priorities/) found that improving employee productivity is now the single biggest business objective driving AI investment, cited by 55% of IT decision-makers. This publication’s own [25th annual State of the CIO research](https://www.cio.com/article/4178006/state-of-the-cio-2026-cios-set-the-course-for-ai-roi.html) tells the same story from the measurement side: lack of clear ROI metrics remains a critical barrier to AI success, cited by 32% of IT leaders, and among organizations that measure AI success at all, operational efficiency and process improvement (40%), employee productivity (34%) and cost reduction (30%) dominate, while revenue impact trails at 27%. And [Deloitte’s State of AI in the Enterprise](https://www.deloitte.com/us/en/what-we-do/capabilities/applied-artificial-intelligence/content/state-of-ai-in-the-enterprise.html) found two-thirds of organizations reporting productivity and efficiency gains from AI, while only 20% can point to revenue growth.\n\nNotice what those numbers describe. The industry didn’t get better at measuring AI. It got better at picking problems that were already measured.\n\nWhich brings us to the diagnostic. When an AI project can’t demonstrate ROI, the instinct is to interrogate the technology. Wrong model. Wrong vendor. Insufficient context. Hallucinations. Sometimes that’s true. But the first question in the post-mortem should be about a decision that was made before a single token was generated: what problem did we pick?\n\nDid that problem have good data behind it? And did it have a scoreboard anyone trusted before the AI showed up? If the answer to either question is no, the project was never going to prove anything, no matter how well the technology performed. You can’t demonstrate improvement against a baseline that doesn’t exist, and you can’t win an argument with a metric that was invented the same week as the pilot. The MIT study’s 95% weren’t all technology failures. A meaningful share of them were selection errors, committed months earlier in a planning meeting, by people who chose an exciting problem over a measurable one.\n\nHere’s the uncomfortable part. Just as the industry figured out the measurability trick, the goalposts started moving.\n\n[Futurum’s survey of 830 enterprise IT decision-makers](https://futurumgroup.com/press-release/enterprise-ai-roi-shifts-as-agentic-priorities-surge/) in the first half of 2026 documents the shift: productivity gains fell from 23.8% to 18.0% as the primary ROI metric buyers use to justify AI investment, while hard financial measures, top-line revenue and bottom-line profitability combined, nearly doubled to 21.7%. The productivity argument carried the pilot era. CFOs accepted “the KPIs moved” as an answer for a while. Now, they want hard dollars.\n\nThis is where the next generation of AI projects will separate winners from the pack, and it requires something almost no one negotiates up front: an ROI exchange rate. That’s the pre-agreed formula, signed off by finance before deployment, that converts KPI movement into currency. One point of first-call resolution improvement equals this many dollars. One hour of engineering time recovered equals that many. It sounds bureaucratic. It’s the opposite. The exchange rate is what lets a project claim its value the moment the KPIs move, instead of spending two quarters in a methodology debate trying to reverse-engineer credit after the fact.\n\nWithout an exchange rate, even a well-instrumented project tops out at a productivity story. With one, the same project is a P&L story. Same technology, same results, entirely different conversation with the CFO.\n\nThis month CIO celebrates the [CIO 100 Awards](https://www.cio.com/), recognizing technology initiatives that deliver measurable business value. Study those winning projects and you’ll find plenty of impressive technology. But the thing they share isn’t a model or an architecture. It’s that “measurable” was engineered in at problem selection. The winners picked problems with real data and trusted scoreboards, and they agreed with finance on what the score was worth before they started playing.\n\nThat’s the part of innovation that never makes it on stage, and it’s the part worth copying. So, flip the question that dominated last fall. Don’t ask where the ROI for AI is. Ask whether you picked a problem that could ever answer that question, and whether anyone wrote down the exchange rate.\n\n**This article is published as part of the Foundry Expert Contributor Network.****Want to join?**", "url": "https://wpnews.pro/news/ais-measurement-crisis-is-over-the-translation-crisis-is-next", "canonical_source": "https://www.cio.com/article/4204035/ais-measurement-crisis-is-over-the-translation-crisis-is-next.html", "published_at": "2026-08-03 12:00:00+00:00", "updated_at": "2026-08-03 12:19:10.004319+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-products", "ai-policy"], "entities": ["MIT", "UC Berkeley", "Foundry", "CIO"], "alternates": {"html": "https://wpnews.pro/news/ais-measurement-crisis-is-over-the-translation-crisis-is-next", "markdown": "https://wpnews.pro/news/ais-measurement-crisis-is-over-the-translation-crisis-is-next.md", "text": "https://wpnews.pro/news/ais-measurement-crisis-is-over-the-translation-crisis-is-next.txt", "jsonld": "https://wpnews.pro/news/ais-measurement-crisis-is-over-the-translation-crisis-is-next.jsonld"}}