{"slug": "the-number-one-metric-for-measuring-roi-post-tokenmaxxing-cost-per-finished-task", "title": "The Number One Metric For Measuring ROI Post Tokenmaxxing: Cost Per Finished Task", "summary": "As enterprises grow cost-conscious about AI spending, the key ROI metric is shifting from raw token usage to cost per finished task, according to Cisco EVP Liz Centoni and LogicMonitor Chief AI Officer Karthik SJ. Centoni said she approves budgets based on cost per finished task rather than token counts, while SJ called token usage a 'vanity metric' and noted a pause on using tokens for their own sake. Gartner research predicts AI inference costs per agentic workflow will increase more than fivefold through 2028, and IDC estimates that by 2028, 70% of top AI-driven enterprises will use advanced multi-tool architectures for dynamic model routing.", "body_md": "# The Number One Metric For Measuring ROI Post Tokenmaxxing: Cost Per Finished Task\n\n## Goals are shifting as leaders become more cost conscious.\n\nTokenmaxxing has dominated the conversation around enterprise tech spending this year, with industry leaders like [Nvidia CEO](https://www.businessinsider.com/jensen-huang-500k-engineers-250k-ai-tokens-nvidia-compute-2026-3) Jensen Huang saying in March that he'd be \"deeply alarmed,\" if a $500,000 engineer didn't consume at least $250,000 in tokens.\n\nAn AI token is a basic unit of text that large language models (LLMs) process, and most AI companies charge their customers based on token usage.\n\nSince Huang's statements in March, the hype has eroded somewhat. [Uber](https://fortune.com/2026/05/26/uber-coo-ai-spending-tokens-claude-code/) blowing through AI budgets in just a matter of months has acted as a cautionary tale for what can happen when token spend isn't closely monitored or tied to ROI.\n\nResearch from [Gartner](https://www.gartner.com/en/newsroom/press-releases/2026-08-17-gartner-predicts-ai-inference-costs-per-agentic-workflow-will-increase-more-than-fivefold-through-2028?utm_campaign=SM_GB_YOY_GTR_SOC_SF1_SM-PR) has found that AI inference costs per agentic workflow will increase more than fivefold through to 2028, with researchers noting that inference cost management has become a top priority for product leaders.\n\nAs business leaders become more cost conscious, there is a new metric guiding decision making; cost per completed task.\n\nLiz Centoni, EVP and chief customer experience officer at Cisco, led the deployment of agentic AI across the company's 20,000-person customer experience organization. Centoni told *International Business Times* that she doesn't look at token spend as a metric and instead prefers to look at cost per finished task.\n\n\"I find that a lot of people talk about tokenmaxxing. If I just looked at the token usage from my team, it doesn't tell me the full picture,\" Centoni said in a video interview. \"I'm approving a budget based on where my teams are looking at the cost per finished task, versus saying, 'oh boy, here are the number of tokens that I'm using.'\" She adds that she spends most of her time talking with her team about picking the right models for the right use case, rather than talking about token cost.\n\n\"I don't think a lot of people are talking about this,\" Centoni said. \"That is maybe one other reason where people budget a certain number and then blow through that number in maybe the first month.\"\n\nKarthik SJ, chief AI officer at cybersecurity provider Logic Monitor, notes that while companies in Silicon Valley emphasized tokenmaxxing to incentivize frontier model usage, he believes token usage itself to be a \"vanity metric.\"\n\n\"I strongly feel like tokenmaxxing doesn't mean token usefulness, because it comes back to fundamentals, like, you know, what exactly are you building and what are you using tokens for,\" SJ said.\n\nSJ notes that by incentivizing tokenmaxxing, there have been instances where employees had been defaulting to using frontier models for tasks like reviewing emails, which he says is not the best use of tokens. \"Now people are more conscious, using the right model for the right task. There's definitely emphasis on, you know, caching, there's emphasis on model routing, that's a piece of the architecture stack that didn't exist,\" SJ said. \"I would say just using tokens for the sake of it, that has definitely come to a pause.\"\n\nFrom this perspective, organizations don't need to use the most expensive frontier models for a given task, with both SJ and Centoni noting the use of model routing to various degrees. More broadly, IDC analysts [estimate](http://www.idc.com/resource-center/blog/the-future-of-ai-is-model-routing) that by 2028, 70% of top AI-driven enterprises will use advanced multi-tool architectures to dynamically and autonomously manage model routing across diverse models to orchestrate complex processes.\n\nIn the case of Logic Monitor, SJ says that the company is looking at how much time each task takes from a human labor perspective, while applying agent workflows to the most high-value, expensive tasks.\n\nMohamed Awad, executive vice president of Arm's cloud AI business unit, an infrastructure provider that just [announced](https://newsroom.arm.com/blog/ibm-and-arm-expanding-ecosystem-for-next-era-of-enterprise-computing) it would be designing a dual purpose chip with IBM Z mainframes, believes cost per token is the wrong metric to track.\n\n\"It's sort of an esoteric metric because it sort of ignores the sort of quality of the token.\" Awad said. \"It's probably going to ultimately end up at, you know, at value delivered or task accomplished, and the sort of cost associated for that task.\"\n\nWith regards to infrastructure costs, Awad said the world \"will move to a place where it's more about the outcome that's delivered as opposed to the underlying piping, and that's going to drive an incredible amount of pressure on optimization of the system.\" In practice, that means everything from land to power, to silicon, racks, warehouses and the model that sits on top of it all, so that providers can deliver answers with the lowest possible cost with the highest fidelity.\n\nThese perspectives indicate that AI spending for its own sake is on the way out. Now business leaders are applying deeper scrutiny to spending, and looking at new ways to measure ROI from adoption.\n\n© Copyright IBTimes 2026. All rights reserved.", "url": "https://wpnews.pro/news/the-number-one-metric-for-measuring-roi-post-tokenmaxxing-cost-per-finished-task", "canonical_source": "https://www.ibtimes.com/number-one-metric-measuring-roi-post-tokenmaxxing-cost-per-finished-task-3807209", "published_at": "2026-09-07 13:38:38+00:00", "updated_at": "2026-09-07 13:55:07.662152+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-infrastructure", "ai-products"], "entities": ["Cisco", "Liz Centoni", "LogicMonitor", "Karthik SJ", "Gartner", "IDC", "Nvidia", "Jensen Huang"], "alternates": {"html": "https://wpnews.pro/news/the-number-one-metric-for-measuring-roi-post-tokenmaxxing-cost-per-finished-task", "markdown": "https://wpnews.pro/news/the-number-one-metric-for-measuring-roi-post-tokenmaxxing-cost-per-finished-task.md", "text": "https://wpnews.pro/news/the-number-one-metric-for-measuring-roi-post-tokenmaxxing-cost-per-finished-task.txt", "jsonld": "https://wpnews.pro/news/the-number-one-metric-for-measuring-roi-post-tokenmaxxing-cost-per-finished-task.jsonld"}}