LAS VEGAS – VMware Cloud Foundation (VCF) Chief Product Officer Paul Turner has a not-so-surprising answer to growing concerns over token costs: private cloud.
“We should stop talking about tokens. Tokens are crazy. Why are we being charged with every little token usage?” Turner rhetorically asked an audience at this week’s VMware Explore event, with everyone also rhetorically acknowledging that token charges are all that’s keeping the AI economy afloat.
Turner’s less cynical view is that organizations should turn to private cloud to better manage surging token costs, and not just because that is a cloud model repeatedly touted by VMware and parent company Broadcom.
“The neat thing about deploying into a private cloud [is] I don't need to worry too much about it,” Turner continued. “I manage infrastructure. I work out what I need in terms of number of GPUs (graphic processing units). I just scale out the GPU resources that I need, and I can consume as (many) tokens as that infrastructure can manage. That's a much nicer way to work on monetization and billing. It's more predictable. I can plan for the future. That's where you need to go. No more tokens.”
While quitting tokens cold turkey may not be feasible, Turner did add that organizations can be smarter in how they dole out these precious resources. One way toward more prudent token usage is in identifying exactly when they are needed. “Not every workload needs a frontier AI model. We've got to stop this,” Turner said. “How many applications do you think need 400-billion parameter models? Not all that many, I'll tell you that.”
Turner said organizations need to look beyond such large language models (LLMs) and instead look at exactly what each application needs and tune their resources toward providing the right data sets to feed each workload.
“You can probably run a chatbot application with a 17-billion parameter model, so look at small language models (SLMs), look at open-weighted models. There’s a lot more out there, including if you look at the ecosystem from different regions of the world. Lots of applications,” Turner said.
This variability was highlighted earlier this year when China-based DeepSeek reportedly slashed its API prices up to 90% amid soaring enterprise token usage.
Token talk #
Token usage and management remain hot topics within the AI ecosystem.
Silicon Data’s tracking of token costs recently sunk below $1 per million tokens, settling earlier this week at just 97-cents per one million tokens. Despite that reported drop, organizations continue to tweak their usage models.
Nutanix CEO Rajiv Ramaswami recently explained that the vendor invested $20 million into building its own AI cluster to help manage token spend tied to internal software development costs, a platform that it’s also now using to underpin some of the vendor’s new agentic AI customer offerings.
Ramaswami during a recent press briefing said the vendor invested that money into its own GPU cluster where it’s now starting to run its open weight models.
“We expect that we will be able to handle a good chunk, up to maybe 80% of our internal needs by this hosted model, hosted on our own clusters, hosting open weight models,” Ramaswami said, adding that “for the remaining we will still go use the best frontier models that are out there.”
Ramaswami explained that Nutanix is using the cluster internally for writing code and coding across the software lifecycle, which includes quality assurance (QA) testing and front-end design. This stack also underpins the vendor’s recently launched Agentic Gateway platform.
“We started out using standard frontier models … Copilot, Cursor, Claude, and the usage has exploded, and so have the costs,” Ramaswami said, adding that the internal AI cluster has allowed the vendor to slash AI-generated software development “on a per-token basis,” and that Nutanix expects the investment to have paid off within one year.
Others are more open to token spend.
AT&T CTO Jeremy Legg told attendees at AMD’s recent Advancing AI event that the carrier was "burning about a trillion-plus tokens per month," a number that's gone up "pretty dramatically."
Legg explained that usage was driven by AT&T’s use of more than 100 generative AI models in production. The firm is spending around 45 billion tokens per day, per slides shown during the keynote, with areas AT&T's models are covering include customer care, fraud detection, and cell tower placement, with Legg even claiming it was "agentifying" things like employee onboarding.