AMD targets tokenomics challenge as enterprise AI deployments scale beyond experimentation
The transition from AI experimentation to full-scale deployment has exposed a stark reality for enterprise customers: AI token routing is emerging as the critical strategy to manage skyrocketing costs as the infrastructure decisions made in phase one come due.
Reactive approaches consistently led to exploding token bills, forcing information technology leaders to urgently rethink their underlying infrastructure investments to achieve sustainable financial returns. Instead of focusing strictly on return-on-investment, IT teams burned through tokens as they scrambled to keep pace with developments in the industry, according to John Hampton (pictured), corporate vice president of global enterprise technical sales at Advanced Micro Devices Inc.
“What we saw happen in phase one of this AI deployment is a ready, fire, aim approach,” Hampton said. “They went out and bought big clusters from our competition and they’re running everything in these frontier models in the cloud – and then they looked at their bill.”
Hampton spoke with theCUBE’s Dave Vellante and John Furrier at the AMD Advancing AI event, during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed the urgent need for infrastructure modernization and the emerging strategies for optimizing token routing in data centers. ( Disclosure below)*
Addressing enterprise AI costs with AI token routing
As organizations push AI beyond simple prompting and into autonomous workflows, the financial burden of running frontier models on large GPU clusters has become a primary concern. Tokenomics — the economics of AI token consumption — has emerged as the defining challenge for enterprise IT leaders navigating this shift, Hampton noted.
“Tokenomics has really become the number one topic for enterprise,” Hampton said. “The affordability and the ROI has really become challenged.”
Instead of relying solely on expensive cloud-based frontier models, companies are adopting a hybrid approach that matches each use case to the most efficient hardware available — whether a CPU or a lower-cost GPU alternative.
“The most optimal solution is not always frontier models with large GPU clusters,” Hampton said. “Many times you can run AI inferencing off CPUs or lower-cost, lower-power GPUs.”
To combat escalating expenses, AMD implemented AI token routing internally, directing workloads to its MI350P GPUs rather than frontier cloud models. The results validated the approach on both cost and performance dimensions, Hampton noted.
“The results of this pilot was a 43% reduction in our token bill,” Hampton said. “And here’s the kicker — we actually increased 2.9x on the response speed.”
Organizations that succeed in the next phase of AI deployment will be those that master AI token routing and the economics of compute alongside hardware modernization. Freeing up power and physical space through data center consolidation gives enterprises capital to reinvest in infrastructure built for AI, Hampton noted.
“My belief is the next generation of AI winners are going to have the best token economies,” Hampton said. “This is what we have to collaborate on as an industry and solve for.”
Here’s the complete video interview, part of SiliconANGLE’s and theCUBE’s coverage of AMD Advancing AI 2026:
( Disclosure: TheCUBE is a paid media partner for the AMD Advancing AI event. Neither AMD, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)*
Photo: SiliconANGLE
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more** 11.4k+ theCUBE alumni**— Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network.
About SiliconANGLE Media
theCUBE AIand theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.