# Agentic AI is breaking the token meter, and enterprises need a plan for what comes next

> Source: <https://siliconangle.com/2026/09/28/agentic-ai-is-breaking-the-token-meter-and-enterprises-need-a-plan-for-what-comes-next/>
> Published: 2026-09-28 20:46:55+00:00

### Agentic AI is breaking the token meter, and enterprises need a plan for what comes next

Per-token pricing was the best thing to happen to enterprises looking to experiment with artificial intelligence, but it may be the worst thing for AI in production.

That’s the quandary at the center of a new [Futurum](https://futurumgroup.com/) report, “[The Off Ramp From Per-Token Pricing](https://futurumgroup.com/research-reports/the-off-ramp-from-per-token-pricing/),” sponsored by neocloud provider [QumulusAI Inc](https://www.qumulusai.com/). The report’s key finding is that agentic AI can drive token consumption per task 10 to 100 times higher than a simple inference call. Agentic AI being more expensive is likely no surprise, but it’s good to see the report quantify it.

This is a real problem I have heard from chief information officers and chief financial officers over and over with increasing frequency. The most successful AI projects end up costing the most, often with budget estimates way off. I spoke with one organization that budgeted $1 million for the year, and the initiative was so successful they spent it in three months.

### Usage pricing punishes success

The appeal of per-token pricing is obvious. A developer can call an application programming interface and have a working prototype by the afternoon, without any capacity planning or procurement cycle. The problem is that the meter doesn’t distinguish between a pilot and a production system serving 20,000 employees, and agents are token machines. A chatbot answers a question. An agent plans, calls tools, checks its work, retries, hands off and summarizes, generating tokens at each step.

Futurum forecasts that agent and reasoning inference will grow by 219% this year, and total inference spending will rise from $120 billion in 2025 to $885 billion by 2030. Put those numbers next to a pricing model that scales linearly with consumption, and the result is a budget line that grows faster than the value it creates.

Mazda Marvasti, co-founder and chief executive of [Amberd.ai](http://amberd.ai/), described the pattern in the report. “When they start deploying it throughout the organization, the cost starts skyrocketing because it’s a useful tool that somebody built, but it’s now priced on a variable basis,” he said. “It starts getting the attention of the CFO and the CIO in terms of how much I’m exactly spending to run this tool, and whether it’s worth it.”

That captures the overlooked part of this story. The risk isn’t just a high bill; it’s that unpredictable bills kill useful projects. Marvasti noted that some customers abandoned internally built automation tools because they couldn’t forecast or justify the costs. That’s a governance failure disguised as a pricing problem, and it will slow AI adoption more than any model limitation. 

### Enterprises have already voted with their capacity

A more interesting data point is that the market has quietly moved beyond the “everything in the public cloud” assumption. According to Futurum’s survey of 824 AI decision-makers, reserved and owned infrastructure account for 66% of AI compute consumption, compared with 19% for on-demand cloud. Some 59% of respondents primarily run AI workloads outside hyperscaler public clouds, in their own data centers, colocation facilities, or with bare-metal providers.

I’d caution against interpreting this as enterprises moving away from the hyperscalers. Much of that owned capacity reflects GPU purchases made when on-demand capacity simply wasn’t available. But it shows enterprises are comfortable making capacity commitments for AI, and the question is no longer whether to commit, but which workloads justify such commitments.

This mirrors the adoption cycle information technology went through with the cloud. Start on demand, discover that steady-state workloads are cheaper on reserved capacity, and end up hybrid. AI is compressing that curve from years to quarters, and agents are the accelerant.

### The real economics are about utilization, not price

One of the more interesting sections of the report is Amberd.ai’s deployment on QumulusAI bare metal. The company partitions an eight-GPU Nvidia H200 server into four virtual environments, each with two GPUs, and tiers customers across them based on latency tolerance.

“With one 8x H200 server, two customers pay for the entire server, and I can probably have about 30 to 35 customers running on that one server,” Marvasti said. “After the second customer, the server is free to me, and any customer that comes after that is profit.” Though those data points are compelling, the lesson isn’t that bare metal is cheap. It’s that Amberd.ai built a custom virtualization layer and tiered pricing to drive utilization. Reserved infrastructure turns a variable cost into a fixed one, and fixed costs only pay off when kept busy. An idle reserved GPU is the most expensive GPU there is.

Futurum’s guidance supports this, recommending reserved bare metal for sustained workloads with predictable utilization above roughly 60%. The report also acknowledges that these environments “require more custom engineering, limiting the operating margin gains for teams without the hardware expertise.”

That caveat deserves more attention than it receives. Most enterprises lack deep bench strength in serving engines, batching, quantization and key-value cache management. Without it, the offramp can lead to a different kind of cost overrun.

### The model question the report doesn’t answer

The offramp works well for open-weight models a company can deploy on infrastructure it controls, which is why Amberd.ai built on private, open-source large language models. But many enterprises have standardized on frontier models available only through their developers’ application programming interfaces or hyperscaler marketplaces. For those workloads, no bare-metal alternative exists, so the pricing lever rests with the model provider.

That makes the reserved-versus-per-token decision a model-strategy decision first. Open-weight models are gaining momentum and are increasingly good enough for classification, extraction, summarization and many agent subtasks. The companies with the most leverage will route work to the cheapest model that does the job well, then run it on the cheapest infrastructure that runs reliably.

In the report, Brennen Smith, chief technology officer of [Runpod](https://www.runpod.io/), explained the importance for finance teams. “If you have a predictable workload, such as a well-defined business operation, a fixed lease contract is the way to go,” he said. “However, if it’s experimentation, scaling, or variable velocity, that’s when you need to go on demand. When talking to CFOs, I recommend budgeting for both.”

It’s the right answer, but a harder organizational change than it sounds, because it requires finance, infrastructure, and AI teams to share a view of workload behavior that most companies lack today.

### What this means for buyers

Per-token pricing isn’t going away, and it shouldn’t. But treating it as the default for production AI is a mistake that agentic workloads will quickly reveal. My advice for IT and finance leaders:

- **Measure cost per task, not cost per token.** Agents change the unit of work. As agent complexity grows, track the cost to resolve a ticket or process a claim end to end.
- **Set a graduation trigger.** Define the utilization and volume thresholds that determine when a workload transitions from APIs to on-demand GPUs and then to reserved capacity. Futurum’s 60% baseline is a reasonable starting point.
- **Be honest about operational skills.** Reserved infrastructure only saves money if it stays busy. If that expertise isn’t in-house, factor in managed services before committing.
- **Test open-weight models now.** Every workload that can run on an open model can move off the meter.
- **Negotiate for hardware cycles.** Contracts should address upgrade paths, renewals, and portability so today’s commitment doesn’t become tomorrow’s stranded capacity.

The companies that win with agentic AI won’t necessarily be the ones with the best models. They’ll be the ones that figured out how to afford to run them at large scale.

*Zeus Kerravala is a principal analyst at ZK Research, a division of Kerravala Consulting. He wrote this article for SiliconANGLE.*

##### Image: [Easy-Peasy.ai](https://media.easy-peasy.ai/b838a505-a75e-435a-89c0-92a8d1257967/f757e2e5-94ad-48d0-a631-a9c427cb92c1.png)

# A message from John Furrier, co-founder of SiliconANGLE:

Support our mission to keep content open and free by engaging with theCUBE community. **Join theCUBE’s Alumni Trust Network**, where technology leaders connect, share intelligence and create opportunities.

- **15M+ viewers of theCUBE videos** , powering conversations across AI, cloud, cybersecurity and more
- **11.4k+ theCUBE alumni** — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network

### Are you an AWS customer?  Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: [https://siliconangle.com/aws-marketplace/](https://siliconangle.com/aws-marketplace/)

##### **About SiliconANGLE Media**

[SiliconANGLE](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fsiliconangle.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=SiliconANGLE&index=9&md5=646b1b564e2259100a2b8638aab0a552),

[theCUBE Network](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecube.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Network&index=10&md5=7de2a85f95ab4a4a495cede20b8cb1da),

[theCUBE Research](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fthecuberesearch.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Research&index=11&md5=7bb33676722925eb57d588ec343e4f6f),

[CUBE365](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.cube365.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=CUBE365&index=12&md5=d310fb35919714e66ad8d42c9c0c1bc6),

[theCUBE AI](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecubeai.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+AI&index=13&md5=b8b98472f8071b23ebb10ab9a8dd0683)and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.
