cd /news/artificial-intelligence/ai-token-demand-is-shattering-foreca… · home topics artificial-intelligence article
[ARTICLE · art-80206] src=io-fund.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

AI Token Demand is Shattering Forecasts

AI token processing demand has shattered forecasts, with annual token consumption already at 135 quadrillion tokens—2.4 times higher than Dell's 2028 estimate of 57 quadrillion. Dell COO Jeffery Clarke said the company revised its 2028 inference token forecast from 1 quadrillion to 57 quadrillion and still expects that number to be too low. Goldman Sachs projects token processing will reach 47 quadrillion per month by 2028, or about 565 quadrillion annually, roughly 10 times Dell's estimate.

read11 min views1 publishedJul 30, 2026
AI Token Demand is Shattering Forecasts
Image: Io-Fund (auto-discovered)

July 30, 2026

Beth Kindig

Lead Tech Analyst

  • Token processing, a key indicator of Inference demand, has surpassed early estimates by leaps and bounds even after large upward revisions.
  • Statements made by one of the market’s leaders in AI server sales illustrates this point clearly.
  • Even as one biggest players in AI inference have seen token processing soar by more than 300X in two years, a top Wall Street bank is calling massive growth ahead.

While Wall Street is busy debating whether AI can monetize, inference demand is exploding, with token processing offering the clearest evidence. Total annual token processing is no longer measured in billions or trillions of tokens, but in the quadrillions and beyond. As annual token processing is now tracked in units with 15 trailing zeros, it is becoming more evident that management teams in the AI supply chain and researchers have drastically underestimated the pace of growth.

Forecast revisions that once seemed dramatic have since been eclipsed years ahead of schedule, with one of the strongest pieces of evidence coming from a top Nvidia partner.

Below, we outline how token processing growth has been greatly underestimated and why expectations may need to move higher once again.

Dell's Token Processing Forecast Miss Underscores the Scale of AI Inference Demand #

Dell is one of Nvidia’s key partners in deploying its GPUs through the company’s PowerEdge servers. Dell’s ability to forecast future demand is critical for supply chain readiness, considering its $16.1 billion in server revenue in Q1 was nearly 3X higher than both HPE and Lenovo and 1.6X higher than Super Micro.

In this context, when we consider the source, a quote from Dell COO Jeffery Clarke is striking. In October 2025, Clarke said: “We thought as we model this, that inference would drive by 2028, 1 quadrillion, that's 15 zeros, 1 quadrillion tokens. Now it's 57 quadrillion, and I'm sure we're wrong.”

mid

In other words, Dell has upped its past estimate for token processing for 2028 by 57X, and still thinks that number is conservative. The company also noted that its expectations for inference demand increased by a minimum of 100X in less than a year.

Based on current levels of token processing, calling Dell’s 57 quadrillion token forecast for 2028 an underestimation is putting it kindly. Tokens processed per day is currently tracking near 370 trillion, or roughly 135 quadrillion per year. This means current token consumption is already 2.4X higher than Dell’s 2028 estimate - with years to spare.

Even at the floor estimate of 300 trillion tokens per day, or 109 quadrillion per year, token consumption today would be 1.9X higher than Dell’s 2028 forecast.

Updated forecasts by industry analysts shed further light on just how far off Dell’s estimate could be.

Analysts and AI Leaders Highlight Token Processing Underestimation #

Goldman Sachs’ forecasts from May 2026 estimate that token processing will hit 47 quadrillion *per month *in 2028. That is approximately 565Q tokens per year, or about 10X Dell’s forecast. The firm sees monthly token processing rising from 1.7Q in mid-2025 to nearly 120Q in mid-2030 (1,440 quadrillion annually) a more than 70X increase in five years. Of this, Goldman estimates that around 101Q tokens will come from agentic workloads, or over 80% of the total.

At a current monthly run rate of 11Q tokens compared to Goldman’s May 2026 estimate of 5.6Q monthly tokens, the bank’s forecast may be conservative.

This chart forecasts rapid growth in global AI token processing, increasing from 1.7 quadrillion monthly tokens in mid-2025 to 47 quadrillion in 2028 and 120 quadrillion by mid-2030. The forecast implies more than 70X growth over five years. The chart also notes a current run rate of 11 quadrillion monthly tokens, roughly double Goldman Sachs' May 2026 estimate of 5.6 quadrillion, and projects that agentic AI will account for approximately 84% of AI workloads by 2030.

Meanwhile, tech consulting firm Tirias Research says when it first forecasted global AI demand in 2023, it expected annual token output would hit 20T by year-end 2024. It estimates that actual token usage hit 667T, more than 33X higher than its original forecast. The company updated its outlook in mid-2025 to 76.9Q tokens annually by 2030, and current token consumption is already around 1.75X higher than this figure.

Another notable data point comes from Anthropic. CEO Dario Amodei said the company planned to achieve 10X growth in 2026. However, in Q1, revenue and usage climbed 80X on an annualized basis, leaving the firm struggling to keep up with its compute needs.

AI Token Processing Is Surging Across Big Tech and Inference Platforms #

While many companies developing AI models have not previously provided token processing expectations to compare against, the raw explosion in token processing would have been hard for anyone to predict.

Google Reports 330X Token Growth in Two Years

If you think that 70X or 33X growth through 2030 is quite a lot, Alphabet’s growth in monthly tokens processed dwarfs that – coming in at 330X over the last two years. Google said that it was processing 3.2 quadrillion tokens per month in May, or 38.4Q a year, which alone would account for 67% of Dell’s 2028 forecast. The growth in the company’s token processing is massive, increasing by 7X since May 2025, and 330X since May 2024.

While these figures measure token processing across all of Google’s surfaces, the company has also seen a huge increase in usage for its Gemini models. In its Q2 2026 earnings call, Google said Gemini models were processing 22 billion tokens per minute, up 120% in just six months, or equal to 1 quadrillion tokens per month. Google's monthly AI token processing increased from 9.7 trillion in May 2024 to over 3.2 quadrillion in May 2026, representing 7X year-over-year growth and accelerating AI inference demand.

Microsoft, OpenRouter, and Fireworks AI Showcase Rapid Token Processing Growth

Microsoft CEO Satya Nadella noted in the company’s April earnings call that it processed over 100 trillion tokens during the quarter, a 5X increase YoY, and processed a record 50 trillion in March. That is a far cry from Google at just 0.1Q tokens per quarter, but shows rapid processing growth nonetheless.

In May, LLM interface provider OpenRouter notes that its weekly token volume increased by 5X in six months from 5 trillion to 25 trillion, and that it was on pace to process more than 1Q tokens in 2026. Additionally, Fireworks AI, which provides an inference serving platform, processes 40 trillion tokens per day as of mid-July, more than doubling in three months from 15T per day in April and up 4X from October 2025 when it hit 10T per day.

Why AI Token Processing Is Growing Faster Than Expected #

Looking at the numbers, the extent to which token processing is far exceeding previous expectations is somewhat staggering. One of the most prevalent reasons for the dramatic increase in token processing compared to what was originally modeled is the rise of agentic AI. As Goldman Sachs puts it plainly; “We weren’t talking about agents a year ago, now we are”.

Anthropic estimates that multi-agent systems use up to 15X more tokens than chatbot requests. Meanwhile, third-party researchers like those at Stanford say that coding agents consume 1,000X more tokens than code reasoning and code chats. Along with this increase in token consumption needs, agent usage is on the rise. Microsoft said in April that its first-party agent usage had increased by 6X year-to-date, or 6X in just four months.

Reasoning Models Take Over Token Consumption

A key enabler of agentic AI is the rise of reasoning models, or models that employ a thought process for how they should respond and refine outputs as they go, introducing significant complexity compared to non-reasoning models. The uptick in reasoning model usage helps explain the vast increase in token processing.

A study that analyzed 100 trillion tokens on OpenRouter found that tokens served by reasoning models increased from 0% at the start of 2025 to around 60% near the end of 2025. The beginning of this trend aligns closely to when OpenAI released the full version of its o1 reasoning model in December 2024.

Researchers also found that prompt tokens per request increased by 4X compared to early 2024, and completion tokens per request tripled. This indicates that users are asking models to complete more complex tasks, increasing the amount of input and output tokens for each request.

The chart tracks reasoning versus non-reasoning token usage on OpenRouter from late 2024 through late 2025. The share of tokens served by reasoning models steadily increases from near zero to more than 60%, crossing 50% by the end of 2025 and highlighting the rapid adoption of reasoning AI models.

Furthermore, Google has shown evidence that queries are rising far faster than raw user count. The company said that from Q2 2025 to Q3 2025, monthly active users on the Gemini app increased by 44% from 450 million to 650 million. However, during the same period, queries increased by 3X, indicating that it not only added many users, but that each user also increased their engagement.

Conclusion #

It is clear that token processing has risen faster than early estimates by multiple orders of magnitude. Even as this has taken place, analysts forecast that explosive token growth will continue for years to come, driven significantly by agentic AI inference.

The continued rise in inference demand may be the strongest driving force behind the AI infrastructure trade going forward. Data center operators will not only require more compute, but also more powerful and efficient computing systems, such as Nvidia’s latest Rubin generation and its inference specific variants.

The I/O Fund recently released our new 90-page Top 20 AI Stocks for Q3 2026 report, where we identify the lesser-known companies best positioned across AI accelerators, memory, networking, energy infrastructure and other critical layers of the AI stack.

Prior Top AI Stock reports have identified **five positions up more than 100% year to date **and ten positions up more than 50% for the I/O Fund, with many held at high allocations. By comparison, the Nasdaq-100 is up just **13% YTD. **

Don’t miss out on the AI trade. Learn more here.

Please note: The I/O Fund conducts research and draws conclusions for the company’s portfolio. We then share that information with our readers and offer real-time trade notifications. This is not a guarantee of a stock’s performance and it is not financial advice. Please consult your personal financial advisor before buying any stock in the companies mentioned in this analysis.

Leo Miller, AI and Semiconductor Investment Writer at I/O Fund, contributed to this analysis.

👉🏻 Share with a Fellow Investor

Help someone else benefit from this insight.

Recommended Reading:

More To Explore

Newsletter

AI Token Demand is Shattering Forecasts

Total annual token processing is no longer measured in billions or trillions of tokens, but in the quadrillions and beyond. As annual token processing is now tracked in units with 15 trailing zeros, i

Nvidia and Google Are Crowding TSMC’s N3 Node - Can Intel Fill the Gap?

Nvidia is moving its next-generation Rubin GPUs from 4nm to 3nm, yet Google’s latest TPUs are already on N3 and are expected to remain there. Meanwhile, a growing number of AI CPUs from Nvidia, Amazon

Intel vs TSMC: How CoWoS Packaging Constraints Could Create an Opportunity for Intel Foundry

Taiwan Semiconductor (TSMC) is the single, most important company to the AI industry. However, to compete with the incumbent, Intel does not need to beat TSMC at leading-edge manufacturing. It only ne

Big Tech’s Free Cash Flow is Turning Negative – Who's Next?

Big Tech’s AI revenue is accelerating, but free cash flow is moving sharply in the opposite direction. Across Google, Microsoft, Meta and Amazon, capex is rising much faster than operating cash flow a

Big Tech Earnings Preview: Is AI Monetization Finally Catching Up to Capex?

The most pronounced difference between 2026’s tech rally compared to rallies in the past is which companies have been left out of it. The names most associated with the AI trade have hardly participat

Nvidia, CXL, and the Battle to Improve AI Inference Economics

This is Part 2 of our two-part series on AI inference economics. In Part 1 — Why Nvidia's Next AI Battle Is About Tokens per Watt, we laid out why tokens per watt has become the defining metric for in

Why Nvidia’s Next AI Battle Is About Tokens per Watt

As hyperscalers move from building AI infrastructure to monetizing it, tokens per watt helps to reflect if revenue is scaling and if profitability is improving. Offload engines can increase tokens per

Micron Is Up 900%. Here’s Why the AI Memory Trade May Still Have Room to Run

Over the past 10 months, memory chip stocks have gone from being solid beneficiaries of the AI boom to capturing a massively outsized piece of the return pie. The inflection in Micron’s performance de

Why the S&P 500 Shrugged Off the Iran War — and What Could Finally Break the Rally

On February 28th, the U.S. went to war with Iran, and the market was handed the kind of shock it hasn't contended with for years. The conflict set off a chain reaction across the region: an ongoing su

Nvidia, CoreWeave, and Nebius: Inside the Circular Financing of the GPU Boom

Neoclouds are one of the more hotly debated AI business models, with CoreWeave and Nebius being the two most widely recognized names. These companies have seen their sales, backlog, and share prices s

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @dell 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-token-demand-is-s…] indexed:0 read:11min 2026-07-30 ·