# Token Growth is Surging - Here Are the Beneficiaries

> Source: <https://io-fund.com/ai-stocks/token-growth-surging-beneficiaries>
> Published: 2026-07-31 00:00:00+00:00

# Token Growth is Surging - Here Are the Beneficiaries

July 31, 2026

### Beth Kindig

#### Lead Tech Analyst

- As token processing has jumped far ahead of industry expectations, compute, networking and power companies are poised to benefit.
- Nvidia Rubin, optical networking, and readily available power are specific solutions that can help data center operators tackle in the huge increase in token processing and inference demand.
- Industry analysts are now forecasting exponential growth in a technology we detailed many months ago.

AI token processing, one of the clearest indicators of inference demand, is blowing past initial expectations. We highlighted in “AI Token Demand is Shattering Forecasts” that Dell raised its 2028 token-processing estimate by 57X, yet actual token processing has already moved far beyond that sharply revised forecast. Google, for example, saw surface-wide token processing rise by 330X from May 2024 to May 2026, while several other players have reported similarly dramatic growth.

Exploding token growth is only one part of the equation. Future GPU generations and model optimizations will continue to improve efficiency and lower cost per token. However, these improvements are driving token consumption higher, creating a cycle where demand growth continues to outpace efficiency gains.

In the second installment of our token processing series, we examine what exploding inference demand means for the AI infrastructure market. While many parts of the AI stack are positioned to benefit, we focus on three specific areas where the implications are significant; compute, networking, and power; highlighting notable companies along the way.

## Compute Efficiency Is Becoming Critical to AI Inference Economics

Inference demand is rising faster than the power available to support it, making token throughput per unit of energy one of the most important metrics in the next phase of the AI infrastructure buildout. As a result, compute systems are being designed to maximize token throughput and token-processing efficiency.

As we detailed in our recent article “** Why Nvidia’s Next AI Battle Is About Tokens per Watt”** increasing tokens per watt is key to hyperscalers growing inference revenue despite power constraints, while also expanding inference margins. Nvidia says Vera Rubin racks will deliver dramatically higher token throughput per GW.

mid

As analysts expect agentic AI to drive the majority of token processing over the coming years, Nvidia says that Vera Rubin is designed to deliver up to 10X more agentic throughput per unit of energy compared to Blackwell. Nvidia reaches this metric by not only increasing raw tokens per second per GW, but also allowing for much higher agent interactivity, resulting in 10X more agents and 2X more tool calls per GW.

Nvidia looks to take inference optimization further through its Vera Rubin + Groq 3 LPX deployments, which it says can deliver up to 35X higher token throughput per MW compared to Blackwell.

For Rubin Ultra, Nvidia has yet to release stats on how it compares to Rubin on token throughput and token per watt, but it will pack 2X more HBM per chip compared to Rubin at 576 GB, and will be available in an NVL576 configuration, connecting eight Rubin Ultra 72 GPU racks. The additional memory and larger NVLink domain should allow models to keep more working memory on high-bandwidth tiers while improving communication efficiency between GPUs. This could increase tokens per unit of energy by reducing the amount of time GPUs spend idle.

## Nvidia's Latest Systems Could Expand Data Center Margins

Data from Morgan Stanley Research supports the idea that moving toward Nvidia’s more advanced systems can deliver significant margin benefits to data center operators. The firm estimates that Blackwell-based data centers process tokens at an approximately 58% net margin. It estimates that this will rise to 78% for Rubin-based data centers, and that Nvidia’s further out Feynman generation will push net margin to 90%.

This comes as processing more tokens within the same power envelope means greater revenue generation at the same level of energy expense. Margin improvements of this, or even close to this magnitude, give data center operators a strong incentive to adopt the latest AI compute systems as token processing soars.

*This chart from Morgan Stanley Research estimates data center net margins from token sales across Nvidia GPU generations. Blackwell-based data centers generate approximately 58% net margins, Rubin-based data centers approach 78%, and Feynman-based data centers reach roughly 90%. The data suggests that higher token throughput and efficiency could significantly improve data center profitability over time.*

## AI Networking Demand Is Accelerating With Agentic AI

Agentic AI and the shift from query-based responses to autonomous agentic workflows is expected to create substantial tailwinds for the networking stack. Increasing tokens consumed per user and per workflow means more data must be exchanged between AI accelerators, CPUs and memory, and between networking fabrics.

Here’s what that means if we look at stats: Arm estimates agentic AI will drive up to a [15X increase in tokens](https://x.com/Arm/status/2044821905815789860) per user and [Nvidia concurs](https://www.nvidia.com/en-us/data-center/lpx/) AgencyBench has forecast that agentic tasks will eventually [require 90 tool calls](https://liner.com/review/agencybench-benchmarking-frontiers-autonomous-agents-in-1mtoken-realworld-contexts), 1 million tokens and will result in hours of execution time.

For a more extreme example, and perhaps one of the biggest case studies yet on the token consumption of agentic AI, [OpenClaw consumed 603 billion tokens](https://www.tomshardware.com/tech-industry/artificial-intelligence/openclaw-creator-burns-through-1-3-million-in-openai-api-tokens-in-a-single-month) from 7.6 million API calls in one month for an estimated $1.3 million in spend. Related to this, Nvidia CEO Jensen Huang estimated in March that combining reasoning models with agents can increase token consumption by roughly 1 million times versus early non-reasoning workloads.

In another supporting metric, Cisco estimates that performing tasks using AI agents increases wide area network (WAN) traffic by 450% compared to humans. This measures data that flows between end users and data centers where inference takes place, rather than directly looking at networking demands within data centers. However, much of this traffic is ultimately still tied to tokens that data center compute must process and output, increasing traffic that flows through networking equipment within data centers. Specifically, Cisco estimates that 70% of this increased traffic comes from AI inference. Agentic AI results in significantly more data entering and exiting data centers, compounding the networking challenge because each task requires more orchestration between GPUs, CPUs and memory.

## Optical Networking Is Emerging as a Critical AI Infrastructure Layer

Within data centers, the requirements are shifting both in scale-out and scale-up domains as data transfer speeds rise and the physical size of AI clusters and pods increase, with optical components becoming a necessity in larger domains as copper hits its physical limits.

Optical transceivers are being used to tackle scale-out requirements as data center operators move to clusters of up to 1 million accelerators, with transceivers and components being a huge growth driver for Lumentum, which saw its sales rise by 90.1% YOY to $808.4 million in its latest quarter.

Co-packaged optics (CPO) is an emerging and longer-term opportunity for Lumentum and other networking players, which will come through both scale-out and scale-up content. One of the key drivers of the scale-up CPO opportunity are pods like the NVL576, where optics helps keep latency at ~320 ns, a 5X improvement to how a similar 576-GPU node could be constructed today, per Corning. Notably, TrendForce is [forecasting explosive growth in the CPO and near-packaged optics](https://www.trendforce.com/presscenter/news/20260615-13098.html) (NPO) market. Overall, it expects the CPO and NPO market to grow from $100 million in 2025 to $39 billion by 2030. This is equal to an astonishing 230% CAGR, or a 390X increase in five years.

Around a month prior to this forecast, we detailed the significant multi-year opportunity in CPO in our article “** Inside Nvidia’s $4B Optical Strategy—and Why CPO Changes Everything”. **We also highlighted

[Lumentum as one of our key winners](https://io-fund.com/ai-stocks/ai-networking-stock-vs-nvidia-7x-returns)in 2026, taking a 9% allocation in January two months before Nvidia invested in the company.

Even as recent jitters around the AI trade have caused the stock to fall more than 30% from its 2026 highs, Lumentum shares remain up over 85% YTD, far ahead of the broader market and tech benchmarks.

For more details on the **I/O Fund’s** **entries** and its diversified **AI portfolio** with **five positions** up **100%+ YTD** and **ten** **up** **50%+**, [sign up here.](https://io-fund.com/premium-services-pricing?utm_source=free_article)

## Data Center Power Demand Is Outpacing Grid Capacity

Finally, increased token processing makes it ever more important for data center operators to secure more power to service inference demand. We [previously noted](https://io-fund.com/ai-stocks/bloom-energy-best-performing-stock-april-2026) that ERCOT’s interconnection queue had surged to approximately 226 GW in mid-November 2025, of which around 165 GW, or 73%, came from data center projects. Over a relatively short amount of time, these figures have increased dramatically.

In mid-June, ERCOT’s large load interconnection queue reached 438 GW, with approximately 390 GW coming from data centers alone. Thus, its data center interconnection queue rose by 136% in just seven months. Meanwhile, ERCOT only expects around [3.9 GW of large load capacity to be energized in Q4 2026](https://www.ercot.com/files/docs/2026/05/24/8-Interconnection-and-Grid-Analysis-Update.pdf). This demonstrates a striking imbalance between data center capacity demand and grid supply, making operators increasingly likely to search for alternatives.

## Power Access Is a Profitability Advantage

This imbalance poses a significant problem for data center operators looking to reap a strong return on their investment. As the Carnegie Endowment for International Peace notes “Countries that can [get data centers online quickly](https://carnegieendowment.org/research/2026/06/the-compute-coalition-how-to-build-the-future-of-ai-in-the-free-world) produce dramatically better returns than those where projects languish in permitting and grid connection queues.”

The organization says that data center project delays have the largest negative impact on life cycle value, with a one-year delay costing a 100 MW data center $500 million, or 5% of its total value. For perspective, nearly 100 GW of data center capacity is expected to be brought online between 2026 and 2030, according to JLL. Extending a one-year delay across all of these projects would lead to $500 billion in added costs.

*This chart from the Carnegie Endowment compares factors affecting U.S. AI data center lifecycle value. Delays have the largest negative impact, with a 1.5-year delay reducing value by 8.9% and a one-year delay reducing value by 5.5%. Changes in power costs, taxes, tariffs, depreciation, and natural gas prices have smaller effects, highlighting timely deployment as a key driver of AI data center returns.*

This makes companies that can provide readily available and behind the meter power to data centers highly valuable partners. These were several of the dynamics that led I/O Fund to designate Bloom Energy as our [Top 2026 Stock pick](https://io-fund.com/ai-stocks/bloom-energy-stock-ai-power). Although shares have tumbled significantly from their 2026 highs, Bloom is still up over 130% YTD, and our initial entry from 2025 is up more than 1,100%, with real-time trade alerts sent to premium members.

## Conclusion

The reality of AI demand growth has shattered early estimates for token processing, yet expectations continue moving up and to the right. This should accelerate demand across compute, networking and power, and other areas of the AI infrastructure trade.

The I/O Fund has used these shifts, along with a close understanding of the key bottlenecks across the AI stack, to inform our investment decisions and generate returns that have outpaced passive tech indexes by a wide margin.

However, Lumentum and Bloom Energy are only two of the lesser-discussed AI stocks our team identified early and positioned in ahead of the broader market. The I/O Fund recently released our **new 90-page Top 20 AI Stocks for Q3 2026 report,** where we identify the companies best positioned across AI accelerators, memory, networking, energy infrastructure and other **critical layers** of the AI stack.

The **I/O Fund** currently has **five positions **up more than** 100% year to date **and **ten positions** up more than **50%**, with many held at high allocations. By comparison, the **Nasdaq-100** is up just **13% YTD. **[Learn more here](https://io-fund.com/premium-services-pricing?utm_source=free_article).

*Please note: The I/O Fund conducts research and draws conclusions for the company’s portfolio. We then share that information with our readers and offer real-time trade notifications. This is not a guarantee of a stock’s performance and it is not financial advice. Please consult your personal financial advisor before buying any stock in the companies mentioned in this analysis. Beth Kindig and the I/O Fund own shares in NVDA, LITE and BE, at the time of writing and may own stocks pictured in the charts.*

*Leo Miller, AI and Semiconductor Investment Writer at I/O Fund, contributed to this analysis. Leo Miller owns shares in NVDA.*

**👉🏻 Share with a Fellow Investor**

*Help someone else benefit from this insight.*

**Recommended Reading:**

### More To Explore

### Newsletter

### Token Growth is Surging - Here Are the Beneficiaries

The reality of AI demand growth has shattered early estimates for token processing, yet expectations continue moving up and to the right. In the second installment of our token processing series, we e

### AI Token Demand is Shattering Forecasts

Total annual token processing is no longer measured in billions or trillions of tokens, but in the quadrillions and beyond. As annual token processing is now tracked in units with 15 trailing zeros, i

### Nvidia and Google Are Crowding TSMC’s N3 Node - Can Intel Fill the Gap?

Nvidia is moving its next-generation Rubin GPUs from 4nm to 3nm, yet Google’s latest TPUs are already on N3 and are expected to remain there. Meanwhile, a growing number of AI CPUs from Nvidia, Amazon

### Intel vs TSMC: How CoWoS Packaging Constraints Could Create an Opportunity for Intel Foundry

Taiwan Semiconductor (TSMC) is the single, most important company to the AI industry. However, to compete with the incumbent, Intel does not need to beat TSMC at leading-edge manufacturing. It only ne

### Big Tech’s Free Cash Flow is Turning Negative – Who's Next?

Big Tech’s AI revenue is accelerating, but free cash flow is moving sharply in the opposite direction. Across Google, Microsoft, Meta and Amazon, capex is rising much faster than operating cash flow a

### Big Tech Earnings Preview: Is AI Monetization Finally Catching Up to Capex?

The most pronounced difference between 2026’s tech rally compared to rallies in the past is which companies have been left out of it. The names most associated with the AI trade have hardly participat

### Nvidia, CXL, and the Battle to Improve AI Inference Economics

This is Part 2 of our two-part series on AI inference economics. In Part 1 — Why Nvidia's Next AI Battle Is About Tokens per Watt, we laid out why tokens per watt has become the defining metric for in

### Why Nvidia’s Next AI Battle Is About Tokens per Watt

As hyperscalers move from building AI infrastructure to monetizing it, tokens per watt helps to reflect if revenue is scaling and if profitability is improving. Offload engines can increase tokens per

### Micron Is Up 900%. Here’s Why the AI Memory Trade May Still Have Room to Run

Over the past 10 months, memory chip stocks have gone from being solid beneficiaries of the AI boom to capturing a massively outsized piece of the return pie. The inflection in Micron’s performance de

### Why the S&P 500 Shrugged Off the Iran War — and What Could Finally Break the Rally

On February 28th, the U.S. went to war with Iran, and the market was handed the kind of shock it hasn't contended with for years. The conflict set off a chain reaction across the region: an ongoing su
