cd /news/ai-infrastructure/cpu-gpu-ratios-and-the-race-to-the-s… · home topics ai-infrastructure article
[ARTICLE · art-136251] src=thediligencestack.com ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

CPU:GPU Ratios and the Race to the Scale Up Domain

The Diligence Stack reported that the CPU-to-GPU ratio for agentic AI inference will vary by hyperscaler, with no standard configuration, and that the scale-up domain is the priority for serving inference at scale. The firm said it raised its datacenter CPU TAM expansion forecast in its May 26 "Secret Agent CPU, Revisited" note after industry conversations strengthened its conviction that dedicated CPU capacity is essential for running agentic systems at scale. The analysis holds that extending the scale-up domain across racks gives operators more tightly connected compute and memory to serve larger models or more simultaneous users economically.

by read11 min views1 publishedSep 21, 2026
CPU:GPU Ratios and the Race to the Scale Up Domain
Image: Thediligencestack (auto-discovered)

Authors note: We are sharing this note, free for all readers. For subscribers these week we have two exclusive interviews we think you will like :)

As the world catches on that agentic AI is going to need a lot more CPUs, both in the enterprise and for personal use cases, the conversation about the CPU-to-GPU ratio is hot again. We have covered this extensively, starting with Secret Agent CPU on March 24, when we looked at how agentic workloads could expand the server CPU market beyond general-purpose servers and head-node CPUs and released one of the earliest TAM expansion forecasts for datacenter CPUs.

In Secret Agent CPU, Revisited, published May 26, we raised that forecast. Our conversations across the industry were strengthening our conviction that dedicated CPU capacity would become an essential part of running agentic systems at scale. We remain bullish on CPUs because of the clarity in inference workload demand and the clear role the CPU plays becoming a larger infrastructure requirement.

But while we talk a lot about the CPU-to-GPU ratio, it is very hard to isolate that number and turn it into a demand forecast. The mix depends on the infrastructure operator and the decisions it makes about serving AI. Key point here, each hyperscalers ratio will be different. There will be no standard CPU/GPU ratio configuration. The main key here is to understand those decisions, we need to understand the scale-up domain and why this is priority number one for serving inference at scale and monetizing compute infrastructure.

Why we add “domain” to scale-up #

A lot gets talked about scale-up, which involves connecting accelerators so they can work closely together in a single compute fabric. We add “domain” firstly because its a helpful mental model to understand networked infrastructure topology and because we want to “connect” how much networked compute and memory can cooperate efficiently on the same workload.

The goal is to make more compute and memory work tightly together as one domain. Today, one way to think about that is a rack whose accelerators cooperate to serve a model to many customers at the same time. Depending on the model and workload, that can mean thousands of simultaneous users, although there is no fixed number of users per rack. Longer conversations require more memory, while more demanding reasoning can occupy the processors for longer. Operators want to expand that domain where doing so lets them serve larger models or more users economically. Across the broader datacenter, they also connect multiple domains, without requiring every processor to participate in one tightly coupled system.

There are two ways this creates demand for more racks. An operator can run additional copies of the model and distribute customers across them. Or a model can require more resources than one rack provides efficiently, which means multiple racks must cooperate on its workload execution.

Extending the scale-up domain across racks gives the operator more tightly connected resources to work with. Models can also run across separate domains through scale-out networking, although the communication cost changes. Connecting racks alone does not establish that they form one scale-up domain.

This is where model size and model context enter the conversation. The model’s weights take up memory, and serving active users requires additional working memory. For many models, the KV cache holds information from the tokens already processed so the model can continue generating its response. Longer context and more simultaneous sessions increase that requirement.

A larger domain gives an operator more room to place the model and its active state. Depending on the workload, that can make a larger model practical to serve or improve throughput at the response time customers expect. The operator still has to balance those benefits: a larger model or longer context can consume capacity that otherwise serves additional users.

And shared memory needs a qualification. Memory remains distributed across processors, with a cost to accessing it remotely.More connected compute does not, by itself, make the same model smarter or more capable. It gives operators more capacity to deliver the capabilities of the models they choose.

Interconnect determines how far the domain can grow #

Our view is that hyperscalers and neoclouds have a strong incentive to expand these domains where the additional scale improves serving economics. Some enterprises will face the same decisions as they deploy larger AI systems on-premises.

The limit is how many accelerators can exchange data fast enough to work efficiently together. When a model runs across multiple accelerators, each needs results from the others to continue its work. If those results arrive too slowly, adding accelerators can leave more expensive hardware waiting for data.

Copper has limits around reach and signal integrity as speeds increase. The underlying boundary depends on the architecture, including its physical layout and power budget.

Optical interconnects expand the possibilities for reach and bandwidth density, which is why we have spent so much time on this part of the market. But the optical system still has to be manufactured, qualified, and serviced. In Optics Won’t Scale as Fast as the Market Expects, we explained why demand for optical connectivity can move faster than qualified supply.

The race is to build larger scale-up domains that can serve larger models and more users concurrently and economically. If the processors spend too much time waiting for data, operators are paying for compute they cannot fully use. How well each company solves that problem will help determine its margins from serving AI at scale.

Where the CPU changes the equation #

The CPUs already present in GPU systems give us a baseline ratio. For example, NVIDIA’s GB200 NVL72 contains 36 Grace CPUs and 72 GPU packages, or one CPU for every two GPU packages. Counting the dies inside those packages would produce a different number, so we need to keep the denominator consistent.

Where the ratio conversation becomes more interesting is when we add dedicated agentic CPU capacity which will manifest itself as dedicated racks of CPUs in the scale up domain of GPU compute racks.

This is why we separate head-node CPUs from dedicated agentic CPUs in our models. While each customer building their own compute racks, like AWS, Azure, and Google, can optimize the number agentic CPUs for their own unique needs the template we have today is of NVIDIA’s Vera CPU Rack which supports up to 256 CPUs and is positioned alongside its accelerator systems. Arm’s AGI CPU also targets agentic infrastructure with AMD, Qualcomm, and Intel following suit.

We do not yet know how much of that capacity every operator will deploy or what their scale up domain mix of CPU/GPU/premium accelerators will be. It could be one CPU rack supporting several GPU racks, or more, with the allocation changing as the workload, or economics, changes.

Consider a useful, but hypothetical, example deployment of a scale up domain of ten GPU racks, each with 72 GPU packages and 36 host CPUs. That gives us 720 GPU packages and 360 CPUs. Adding a dedicated rack containing 256 CPUs takes the total to 616 CPUs, or about 0.86 CPUs per GPU package. Two dedicated CPU racks take it to 872 CPUs, or about 1.21 to one.

That is math using stated configuration assumptions, not a deployed-system claim or a forecast. We highlight this use case to show how varied potential configurations can be and why trying to estimate any CPU:GPU ratio is not the right analysis or measurement of the market.

Demand still has to fit what can be manufactured #

Our CPU forecast (available in institutional tier of CS Atlas) separates general-purpose servers, accelerator head nodes, and standalone agentic systems along with key vendor share assumptions. We then constrain shipments using our manufacturing assumptions since our forecasts are grounded also in what’s manufacturable.

Our most current base case implies the following growth from a 2026 baseline:

  • 2027: CPU silicon-value growth of29% ; CPU package shipment growth of20% .
  • 2028: CPU silicon-value growth of33% ; CPU package shipment growth of22% .
  • 2029: CPU silicon-value growth of25% ; CPU package shipment growth of11% .
  • 2030: CPU silicon-value growth of35% ; CPU package shipment growth of29% .

These are conditional, and continually updated, Creative Strategies estimates, grounded in our most current assumptions and industry checks. Silicon value includes an imputed value for captive CPUs, so it is broader than merchant CPU revenue. Pricing and product mix explain why dollar growth can exceed shipment growth given a core assumption we are making on ASP for agentic CPUs specifically.

The model constrains CPU output by compatible foundry-node capacity and allocation after competing demand, then applies yield, assembly, and substrate assumptions. Its 2029 slowdown and 2030 acceleration depend heavily on node transitions and the capacity those transitions make available. For customer/clients with access to our data models we break down the full assumptions in base/bull/stretch scenarios.

We remain aggressive in our forecasts but also grounded in manufacturing reality. Wafer mix assumptions matter and we cannot assume every wafer is available to CPUs while GPUs and custom accelerators are competing for capacity. Nor does the model establish the industry’s absolute manufacturing ceiling. We will continue refining these assumptions as new information becomes available on customer allocations and system-level deployment constraints.

Any ratio estimate has to be grounded. We would therefore be careful about declaring that a two-to-one CPU-to-GPU ratio, or more, is either inevitable or impossible. It first needs a defined workload and counting convention, then a test against the supply model aligned with tangible tracking of hyper scaler scale up domain mix. Demand above our forecast and demand above feasible supply are different conclusions.

Each operator is making a hardware bet #

Hyperscalers make distinct decisions about the infrastructure they want to build, working with system suppliers and manufacturing partners. They choose the balance of merchant and custom silicon, then determine the memory and networking required to make those processors productive within a TCO profile.

It is important to note how one operator may prioritize flexibility across many customer models. Another may have enough predictable internal demand to justify a more specialized accelerator. They can also assign different stages of inference to different systems when the benefit outweighs the cost of moving state between them. Every implementation is unique and many times bespoke which is why generalities harm more than hurt the analysis.

Mixing merchant and custom infrastructure across a fleet is different from making those processors cooperate inside one tightly connected domain. That requires compatible hardware and software. These decisions commit capital to assumptions about workloads that are still evolving, and those assumptions help determine the margin profile of serving AI at scale.

Our conviction is that enterprise and personal agents will create much more work around the model layer in the scale up domain. We will have many simultaneous sessions calling on CPUs and then returning to accelerators, with memory retaining the state needed to continue. The amount of dedicated new CPU infrastructure depends on how much existing capacity absorbs and how efficiently the software schedules that work. Another reason, we are bullish on the growth profile of AWS, GCP, and Azure.

For who benefits here in silicon, we have high conviction on both AMD and Intel, with Intel standing to doubly benefit with Intel Foundry well positioned to also manufacture CPUs as additional capacity is brought online. We maintain NVIDIA is also a primary beneficiary with their agentic CPU roadmap, and Arm and Qualcomm also coming online with solutions. This is not the server CPU TAM of old where share taking was the fundamental analytical baseline. This is a significant TAM expansion and the analysis now who is are the unit volume and price share gainers in a market that is growing in dollar value ~30% YoY in our supply constrained forecast. What would make us less bullish is evidence that CPU work per completed task falls enough to offset usage growth, or that existing infrastructure absorbs much more demand than we expect. For now, we believe the expansion of agentic workloads supports more CPU capacity, with operators making different decisions about how much belongs around each accelerator deployment.

That is how we think about the CPU-to-GPU relationship: through the complete inference system and how its scale-up domains shape what operators can manufacture and deploy economically. Keeping that view grounded is our ongoing research goal as hardware choices evolve and we learn more about how these systems perform in practice.

Continue the research in Atlas #

Subscribers with CS Atlas access can use Atlas to connect our CPU reports with the research on networking and optical supply. Ask how dedicated agentic CPU racks change the processor mix, which assumptions constrain growth, and what evidence would change the forecast.

Explore this research in Atlas

The link opens an editable question after sign-in. Answers depend on the research and model versions available with your access.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @the diligence stack 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/cpu-gpu-ratios-and-t…] indexed:0 read:11min 2026-09-21 ·