As agentic AI inference surges, tokenomics becomes the enterprise’s defining budget constraint
The transition from chatbots to autonomous agents is changing the shape of demand itself, and tokenomics — the economics of AI token consumption — is emerging as the defining constraint on enterprise budgets as round-the-clock inference replaces intermittent usage. Fewer than 1% of potential users are currently deploying agents at scale, leaving enormous headroom for growth in inferencing demand, according to Jeetu Patel (pictured), president and chief product officer of Cisco Systems Inc.
“There’s less than 1% of the world that’s using agents,” Patel said. “If you believe that agents are actually a step function improvement from a chatbot, where you can actually have either your personal productivity or your company’s productivity materially change as a result of agents, I don’t see how you don’t stay in a continued kind of supply shortage for a pretty long time period.”
Patel spoke with theCUBE’s Dave Vellante at the AMD Advancing AI event, during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed distributed inference, tokenomics and the shifting economics of agentic AI deployment. ( Disclosure below.)*
Tokenomics emerges as the defining constraint on agentic AI
Unlike chatbots, agents don’t wait for a human prompt — they run continuously and increasingly talk to other agents, which is pushing a new class of machine dedicated purely to agent workloads, Patel explained. Cisco’s answer is Cisco Cloud Control, a unified management plane built to track that sprawl across every location inference can occur.
“It can go out and look at all of your inferencing capacity that you have in the cloud, all of the inferencing capacity you might have in a private data center, as well as your endpoint devices, in a single management plane,” Patel said. “Within that hybrid inference model, AMD handles intelligent routing and compute while Cisco supplies network bandwidth, security and what Patel called tokenomics — visibility into which agents are consuming tokens and whether that usage has become excessive.”
Enterprises are also being pushed toward smaller, task-specific models to keep costs in check rather than defaulting to frontier systems for every job. Cisco recently released Antares, a family of open-weight models built specifically to locate vulnerabilities in code without sending proprietary data to the cloud, a direct example of that routing logic in practice, Patel noted.
“That particular announcement doesn’t mean that we actually don’t partner very deeply with Anthropic and with OpenAI,” Patel said. “It just means that you’re going to need to have [small language models] that you want to have control over. If I can find 70% of the vulnerabilities that way, great. And the remaining 30% I can find with a frontier model.”
Ease of use, not compute, remains the biggest barrier to broader agent adoption. The technology still has to travel a long distance before it reaches ordinary users the way the internet once did, Patel noted.
“We’re moving at a pretty fast pace, but we’re nowhere near the level of simplicity that’s needed for eight billion people all going out and activating thousands of agents,” he said.
Here’s the complete video interview, part of SiliconANGLE’s and theCUBE’s coverage of the AMD Advancing AI event:
( Disclosure: TheCUBE is a paid media partner for the AMD Advancing AI event. Neither AMD, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)*
Photo: SiliconANGLE
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more** 11.4k+ theCUBE alumni**— Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network.
About SiliconANGLE Media
theCUBE AIand theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.