# The Agent Measurement Problem: Five Competing Metrics, No Standard

> Source: <https://forkast.news/the-agent-measurement-problem-five-competing-metrics-no-standard/>
> Published: 2026-09-11 11:28:33+00:00

Salesforce is currently touting 7 billion “Agentic Work Units”—its proprietary measure for tasks like updating records or closing cases—as the ultimate proof of momentum for its Agentforce platform, which has hit $1.5 billion in annual recurring revenue. But look past the headline numbers, and you find a jarring disconnect. While Salesforce celebrates this massive surge in activity, broader enterprise data suggests that companies are struggling to translate that volume into anything resembling a stable production environment.

We are currently trapped in a measurement vacuum. There is no standard for what success looks like in agentic AI, so we are left with five competing metrics that measure entirely different things. Salesforce tracks activity volume through its Agentic Work Units. Gartner warns of a looming cliff, predicting that over 40% of agentic projects will be canceled by 2027 due to costs and unclear value. McKinsey highlights a budget crisis, finding that 93% of enterprises are overspending on AI. The [Futurum Group](https://futurumgroup.com/press-release/enterprise-ai-roi-shifts-as-agentic-priorities-surge/) tracks a shift in ROI, noting that decision-makers are moving away from productivity metrics toward direct financial impact. Finally, S&P Global Market Intelligence and MIT data show a stark gap between the 80% of apps embedding AI and the mere 31% of organizations actually running agents in production.

Here is the thing: these numbers do not actually conflict. They simply measure different stages of a messy lifecycle. The problem is that enterprises are trying to manage a multi-million dollar transition using a fragmented dashboard. As Anushree Verma, a senior director analyst at Gartner, notes, most of these projects are still early-stage experiments driven by hype. This immaturity is actively distorting how companies spend their capital. Furthermore, the [MIT NANDA report](https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/) highlights that 95% of generative AI pilots are failing to meet CFO expectations, underscoring the disconnect between initial excitement and actual business results.

Consider the current misalignment in budget allocation. While the most measurable ROI for AI agents is consistently found in back-office automation, more than 50% of GenAI budgets are still being funneled into sales and marketing. Companies are chasing the growth narrative while ignoring the operational efficiency that actually pays for the technology. This is compounded by a hidden cost structure: McKinsey research shows that 60% of total agentic AI spend is consumed by iterative response refinement—the cycles of checking, correcting, and improving outputs—rather than the initial inference that generates the value.

Keith Kirkpatrick of the Futurum Group explains that while the productivity argument was the right metric for the pilot phase, the market has matured. Enterprises are now demanding that every AI capability connect directly to revenue growth or margin improvement. We need to move toward standardized, cost-per-outcome metrics that account for the full lifecycle of an agent, not just the number of tasks it triggers.

With [Dreamforce 2026](https://uctoday.com/dreamforce-2026-5-productivity-automation-developments-to-watch) approaching this September, Salesforce will undoubtedly showcase its impressive activity numbers. For enterprise decision-makers, the challenge is to look past the volume. Instead of asking how many work units an agent can process, buyers should ask harder questions: What is the cost of the refinement cycles required to keep that agent accurate? How much of the budget is tied to actual margin improvement versus experimental overhead? And most importantly, how does the agent’s performance hold up when it moves from a controlled pilot to the unpredictable reality of production? If the answer is just more tokens, the project is likely a liability, not an asset.
