cd /news/ai-infrastructure/the-physical-stack-behind-ai Β· home β€Ί topics β€Ί ai-infrastructure β€Ί article
[ARTICLE Β· art-142939] src=intelligence.mts.now β†— pub= topic=ai-infrastructure verified=true sentiment=Β· neutral

The Physical Stack Behind AI

A new attributed record of AI compute infrastructure catalogs 83 facilities, 482 GPU clusters and 2,935 price observations, finding that official H100 rental prices range from $4.57 to $18.37 per accelerator-hour across 155 catalog observations β€” a 4.0Γ— spread within a single hardware generation. The record estimates the largest facility, Colossus 2, at 946 MW, and reports that the four largest infrastructure spenders logged $425.2B in combined capital expenditures in their latest fiscal years against $130.7B of depreciation. NVIDIA has proposed bringing 800-volt direct current closer to the rack, a design the record says would cut conversion stages and waste less energy, though most operating data centers do not use that architecture today.

read35 min views1 publishedOct 1, 2026

An attributed record of where AI compute sits and what it costs: facilities, GPU clusters, power, the chips themselves, cloud prices, measured training performance, and the companies building each layer.

What the record shows #

  • The largest facility in the record, Colossus 2, is estimated at 946 MW . All twelve of the largest facility figures are estimates (Figure 2).
  • Dated construction records show how quickly the biggest campuses can grow, in some cases reaching hundreds of megawatts of IT power within roughly a year, with one announced plan extending to 1,925 MW by 2028 (Figure 1).
  • Official prices vary widely even for the same accelerator. H100 rates run from $4.57 to $18.37 per accelerator-hour across 155 catalog observations, a 4.0Γ— spread within one hardware generation (Figure 4).
  • The four largest infrastructure spenders reported $425.2B of combined capital expenditures in their latest fiscal years, compared with $130.7B of depreciation. Much of today's spending will reach the income statement over the useful lives of those assets (Figure 5).

The underlying record contains 83 facilities, 482 GPU clusters, 2,935 price observations, and SEC facts for 53 public companies. Every value is reported, estimated, or derived, with its source attached.

How the system fits together #

Jensen Huang calls the AI economy a five-layer cake: energy, chips, infrastructure, models, and applications. NVIDIA calls the physical system within that infrastructure layer an AI factory: a data center designed to turn data and electricity into intelligence. A site takes shape in the order land β†’ shell β†’ power β†’ cooling β†’ compute. The sections below follow that build, then move inside the building from chips to racks and clusters, before looking at virtual machines, workloads, and economics.

Build the site

An AI campus is built in sequence: land, shell, power, cooling, then compute. We begin with the grid because the available power often sets the pace and ultimate size of the project.

01Grid + power

A data center starts with a power connection. Generation, transmission, substations, and backup systems determine how much electricity the campus can use. A headline power figure might describe capacity that is contracted, permitted, energized, or already available to servers. Each marks a different stage of the project.

NVIDIA has proposed bringing 800-volt direct current closer to the rack. The design would require fewer conversion stages, waste less energy, and leave more room for computing equipment. It describes a possible transition from today's AC facilities through hybrid systems to native 800 VDC sites. Most operating data centers do not use this architecture today. NVIDIA 800 VDC architecture β†—

What's inside: 4 components, their suppliers and sources #

Grid supply Generation

The generation fleet and wholesale market that supply energy to the local utility or balancing authority. Nameplate generation is not the same as firm capacity available to a continuously loaded AI campus.

What to watch. Is the power physically deliverable and firm through peak conditions, or is the announcement only an energy purchase or aspirational generation build?

Firming + backup Reliability

Batteries, on-site generation, utility reserves, and backup systems bridge outages and variable supply. Backup capacity may be permitted for emergencies without being authorized for continuous operation.

What to watch. How many hours can the system carry the IT load, what emissions or operating limits apply, and does the backup design support uptime without becoming stranded capex?

Interconnection + substation Delivery

Transmission upgrades, transformers, switchgear, and the interconnection agreement convert a grid promise into power that can reach the campus. This is often the schedule-critical path.

What to watch. Is the project in a queue, under study, contracted, under construction, or energized, and who pays for network upgrades if scope or timing changes?

Delivered campus power Usable capacity

Gross utility service is reduced by power conversion, cooling, and other facility loads before it becomes IT power available to servers. The conversion ratio is usually summarized by power usage effectiveness.

What to watch. How much contracted capacity is energized, how much reaches IT equipment, and how quickly can racks consume it without waiting for cooling or network completion?

Market participants9 companies #

Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.

Company Role in this layer Revenue Capex PP&E, net
BE On-site fuel-cell generation $746.4M $26.2M $401.1M
CAT On-site generation and backup power $20.5B $1.3B $15.6B
CEG Firm and nuclear generation $5.8B $2.5B $41.2B
CMI Backup generators and distributed power $9.5B $438.0M $7.0B
DUK Regulated utility and grid delivery $9.0B $4.1B $130.0B
ETN Switchgear and power distribution $8.5B $446.0M $4.7B
POWL Electrical distribution and control equipment $311.7M $10.4M $118.6M
PWR Transmission and substation construction $9.6B $451.0M $3.7B
VRT UPS, power delivery, and cooling $3.3B $285.9M $1.2B

02Facility

A data-center campus can contain several data halls, the large secured rooms where rows of racks are installed. The halls open in phases as their power, cooling, security, and network systems are commissioned. Almost every watt consumed by computing equipment becomes heat, so the cooling system limits how many racks each hall can support and how long they can run at full power.

The twelve largest facility records, each unit square 25 MW of estimated facility power. All twelve are estimates, and several describe campuses still under construction; the largest, Colossus 2 at 946 MW, is roughly the electrical draw of a mid-sized city.

What's inside: 4 components, their suppliers and sources #

Substation + electrical yard Power conversion

High-voltage service, transformers, switchgear, UPS systems, and distribution equipment step grid power down and route it safely to data halls.

What to watch. Transformer and switchgear lead times can gate energization even after utility capacity is awarded. Track redundancy and the difference between ordered and installed equipment.

Cooling plant Thermal infrastructure

Chillers, cooling towers, heat exchangers, pumps, and liquid distribution remove heat from dense racks. Rack architecture determines how much heat must be rejected to air versus liquid.

What to watch. Does the cooling design support the rack density being purchased, and are water, heat-rejection, and mechanical permits aligned with the compute delivery schedule?

Data halls Deployable floor

Secured white space, busways, cooling distribution, and network pathways where racks are installed in phases. A campus announcement can include future halls that are not yet commissioned.

What to watch. How many halls are shell-complete, powered, commissioned, and occupied, and which reported capex belongs to the current phase versus the full master plan?

Control + network rooms Operations

Building-management systems, security, telemetry, carrier rooms, and operations tooling keep power, cooling, and networks observable and available.

What to watch. Physical completion is not the same as operational readiness. Commissioning, carrier diversity, monitoring, and trained operations staff determine when revenue-producing workloads can begin.

Market participants12 companies #

Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.

Company Role in this layer Revenue Capex PP&E, net
ACM Engineering and program management $3.6B $99.9M $459.7M
CARR Chillers and HVAC systems $5.3B $94.0M $3.1B
FIX Mechanical and electrical construction $3.3B $288.8M $653.9M
ETN Electrical distribution and protection $8.5B $446.0M $4.7B
EME Electrical and mechanical construction $5.2B $59.9M $278.6M
J Data-center engineering and delivery $4.1B $61.7M $311.6M
JCI Cooling and building controls $6.1B $148.0M $2.1B
MOD Data-center cooling systems $874.1M $46.4M $536.1M
NVT Enclosures and electrical protection $1.5B $57.6M $447.9M
PWR Electrical infrastructure construction $9.6B $451.0M $3.7B
TT Chillers and thermal management $6.4B $156.1M $2.4B
VRT Power and thermal infrastructure $3.3B $285.9M $1.2B

Assemble the compute system

Compute is assembled from the inside out. Chips sit in rack-scale systems with memory, networking, power, and cooling. Multiple racks are then linked into a cluster.

03Chips

Inside each rack, GPUs and CPUs do different jobs. GPUs handle the model's parallel math, and inference throughput is usually measured in tokens. CPUs handle sequential work such as tool calls, code execution, and browser sessions. An agent can move between the two hundreds of times before it finishes a task.

What the chips measure

GPU throughput is usually measured in tokens. For agentic systems, CPU performance is easier to understand in completed tasks and concurrent agents. The denominator then tells you whether the claim is about energy, cost, serving demand, or the output of an entire site.

The output of a GPU is tokens, produced by parallel math on thousands of cores. The output of a CPU is tasks: the sequential steps around the model, such as tool calls, code, and browsers.

A token count only becomes a claim once it has a denominator. Tokens per watt is a question about energy: how much model output remains after power conversion, cooling, and the chip's own efficiency are taken into account, which is the figure the power layer sets (01 Grid + power). Tokens per dollar is the commercial question: tokens divided by the meter you actually pay, whether a VM-hour, an accelerator-hour, or the depreciation on hardware you own, and the same chip gives different answers under different meters (06 Virtual machine). Tokens per user is throughput per person served: how many concurrent users one system holds at an acceptable speed, where batching, the KV cache, and latency targets all trade against each other (the inference process). Tokens per AI factory treats the whole plant as one machine: NVIDIA's unit for a gigawatt-scale campus designed and operated as a single system, with its DSX blueprint as the reference design for one (02 Facility).

The CPU's units are tasks per second and agents per CPU. Token throughput describes the model running on the GPU. Task throughput describes the sequential work surrounding it. Public benchmarks for complete agent workflows are only beginning to emerge, so this page does not assign the CPU a performance figure.

What's inside: 4 components, their suppliers and sources #

GPU + HBM Parallel compute

Thousands of cores execute the same operation on different data at once, with high-bandwidth memory stacked beside the die so the math never waits on the wires. Its output is tokens.

What to watch. A tokens-per claim needs its denominator stated: per watt is a power question, per dollar is a meter question, per user is a serving question, per AI factory is a site question.

Host and agent CPU Sequential compute

Fewer, faster cores for work that has to happen in order: tool calls, code execution, browsers, sandboxes, data pipelines, and orchestration beyond the model. NVIDIA's Vera CPU is purpose-built for this agentic work, with 88 custom Olympus cores and LPDDR5X memory. Its output is tasks.

What to watch. Ask what percentage of an agentic workload's time runs on the CPU, and whether the host CPU is sized so the GPU is not idling at full price while a tool call finishes.

The agent loop GPU-to-CPU link The model runs on the GPU, hands a step to the CPU, and the CPU runs it and goes back to the GPU to ask what's next. Agents work on a branchy decision tree, so one task can cross this link hundreds of times. A slow CPU step leaves the expensive GPU waiting.

What to watch. Compare GPU-to-CPU ratios across systems and VM shapes; the ratio was designed before agentic workloads arrived, and agent-workflow benchmarks that would validate it do not exist yet.

Units of output Measurement

NVIDIA measures GPU output in tokens per watt, per dollar, per user, and per AI factory, where an AI factory is a whole gigawatt-scale site designed as one machine. CPU output is tasks per second and agents supported per CPU.

What to watch. Two claims with the same numerator can be answering different questions. Restate every throughput number with its denominator before comparing it to anything on this page.

Market participants7 companies #

Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.

Company Role in this layer Revenue Capex PP&E, net
GOOGL TPU accelerators and Axion CPUs $119.8B $80.6B $321.2B
AMZN Trainium accelerators and Graviton CPUs $200.6B $173.0B $446.0B
AMD Instinct GPUs and EPYC host CPUs $11.5B $1.2B $3.4B
AVGO Custom accelerators for hyperscalers $22.2B $481.0M $2.8B
MRVL Custom accelerators and interconnect silicon $2.4B $155.7M $972.5M
MU High-bandwidth memory and server DRAM $41.5B $19.6B $56.4B
NVDA GPUs, Grace and Vera CPUs, NVLink $81.6B $1.8B $12.4B

04Rack

A modern AI rack contains far more than GPUs. The reference system shown here combines CPUs, GPUs and their HBM, networking, local storage, management hardware, power shelves, busbars, and liquid cooling. Every part affects how the rack performs, what it costs, and how quickly it can be installed.

The rack is increasingly the unit of competition. CPU, GPU, HBM, networking, storage, power, management, and liquid cooling all affect the performance of the system. NVIDIA's DSX blueprint extends that co-design to the whole AI factory, connecting reference systems, simulation, operations software, facilities guidance, and partner technologies. NVIDIA defines and validates the reference architecture, while its partners build and operate the sites.

What's inside: 4 components, their suppliers and sources #

NVLink switch system Scale-up fabric

Nine 1RU switch trays connect the 72-GPU NVLink domain through a passive copper cable backplane. Each tray contains two NVSwitches with 72 NVLink ports plus its own provisioning, telemetry, security, and control hardware.

What to watch. Scale-up networking is part of the rack bill of materials and system yield. It is not captured by multiplying a standalone GPU price by 72.

Compute trays Compute + host system

Each of the 18 liquid-cooled 1RU compute trays contains two Grace CPUs and four Blackwell GPUs, plus cluster networking, BlueField DPUs, local NVMe, management controllers, and the operating-system image.

What to watch. The deployable unit carries substantially more content than accelerators alone: CPUs, DPUs, NICs, storage, boards, cold plates, management, assembly, and support all affect price and lead time.

HBM3e GPU memory High-bandwidth memory

HBM sits in the GPU package and supplies model weights and activations at far higher bandwidth than ordinary server memory. NVIDIA specifies the aggregate rack capacity and bandwidth; it does not identify the memory supplier in this product specification.

What to watch. Track HBM capacity per accelerator, stack generation, bandwidth, packaging yield, and supplier qualification. HBM availability can constrain GPU shipments and shift value toward memory suppliers.

Power + liquid cooling Rack infrastructure

Eight power shelves convert AC input to nominal 50–51V DC and distribute it over a busbar with N+N redundancy. Liquid manifolds and cold plates cool CPUs and GPUs while other components remain air cooled.

What to watch. A roughly 120kW rack changes the facility bill of materials. Delivery is constrained by electrical distribution, liquid loops, commissioning, leak detection, and the facility’s ability to accept dense racks.

Market participants9 companies #

Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.

Company Role in this layer Revenue Capex PP&E, net
AMD Accelerators and rack-scale systems $11.5B $1.2B $3.4B
APH High-speed and power interconnects $8.8B $647.1M $2.9B
DELL Enterprise AI systems and racks $43.8B $963.0M $6.9B
HPE AI systems and liquid-cooled racks $10.7B $1.2B $5.6B
MU High-bandwidth memory $41.5B $19.6B $56.4B
NVDA Reference racks, GPUs, CPUs, and NVLink $81.6B $1.8B $12.4B
SMCI Server and rack integration $10.2B $133.8M $607.7M
TEL Power and data connectivity $5.2B $832.0M $4.5B
VRT Rack power and liquid cooling $3.3B $285.9M $1.2B

05Cluster

A cluster links many racks into one computing system. High-speed networking keeps the racks in sync, shared storage feeds them data, and scheduling software assigns the work. The cluster is operational only when all of those pieces have been commissioned and workloads can run reliably.

Zooming back out from one machine to all of them: every cluster record with a first-operational date and a scale estimate, 2016 to present. The march up the log axis is the buildout.

What's inside: 4 components, their suppliers and sources #

Scale-out fabric Inter-rack networking

Ethernet or InfiniBand switches, adapters, optics, and cables connect racks into a training system. Fabric topology determines how efficiently additional accelerators contribute to a distributed job.

What to watch. Does networking capex and optical content rise faster than accelerator count, and does the delivered topology provide enough non-blocking bandwidth for the target workloads?

Compute racks Installed capacity

Configured accelerator racks supply the physical compute. Physical chip count, rack count, commissioned capacity, and H100-equivalent capacity are different measurements.

What to watch. Distinguish ordered, delivered, installed, networked, and operational racks. Usable capacity depends on the system around the chips and the power state of the facility.

Shared storage Data plane

Parallel filesystems and object storage feed training data, absorb checkpoints, and recover jobs. Slow checkpoint or input pipelines can leave the accelerator fleet idle.

What to watch. Can the storage layer sustain workload throughput and failure recovery at cluster scale, and is storage or data movement becoming a material share of cost per useful token?

Control plane Orchestration

Schedulers, health monitoring, provisioning, and failure recovery turn a hardware fleet into a service that model teams can use continuously.

What to watch. How much installed capacity is actually available and productively scheduled, and what software or reliability advantage lets one operator earn more from the same hardware?

Market participants14 companies #

Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.

Company Role in this layer Revenue Capex PP&E, net
ANET AI Ethernet systems $3.0B $84.2M $312.5M
ALAB PCIe and CXL connectivity $392.4M $28.1M $119.3M
AVGO Ethernet switch silicon and custom accelerators $22.2B $481.0M $2.8B
CLS Networking and compute systems manufacturing $4.7B $493.3M $1.0B
CSCO Networking, optics, and systems $15.8B $1.0B $2.6B
COHR Optical components and transceivers $7.1B $1.1B $3.0B
CRWV GPU-cluster operator $2.6B $14.1B $46.7B
CRDO High-speed connectivity and DSPs $1.3B $57.3M $101.6M
FN Optical and electronics manufacturing $4.6B $252.5M $615.1M
LITE Optical components for data-center links $3.0B $451.3M $1.2B
MRVL Interconnect, optics, and custom silicon $2.4B $155.7M $972.5M
NTAP Enterprise data and storage systems $6.9B $198.0M $592.0M
NVDA Accelerators and scale-up fabric $81.6B $1.8B $12.4B
PSTG Training data and checkpoint storage $1.1B $68.4M $613.9M

Put capacity to work

Once the cluster is commissioned, customers can rent the capacity and operators can measure what it produces. Workloads and utilization then determine whether all that spending produces a return.

06Virtual machine

Cloud customers rent this hardware through virtual machines. Cloud providers expose this hardware as named instance types and charge by time. The price usually covers GPUs, host CPUs, system memory, local storage, networking, and platform services, so two hourly rates may include very different amounts of hardware.

What cloud capacity costs: 2,935 price observations from official cloud catalogs, restricted to on-demand meters with a disclosed accelerator count so rates compare per accelerator-hour. Each hardware generation enters the catalog at a higher rate, and the same accelerator spans a wide range across regions and VM shapes, a 4.0Γ— spread for the H100. A catalog rate is a public list price, not the average price customers actually pay.

What happens when a model serves a request

The model's weights must be loaded into accelerator memory before inference can begin. A request then passes through the same sequence for every token it generates. This explains what the machine is doing; the workload section below starts with traffic and estimates the fleet required to serve it.

The trained parameters are the model. They begin as files on storage, and before inference can run those values must fit in accelerator memory at the chosen numerical precision. HBM capacity therefore sets the minimum hardware footprint: large models are sharded across multiple accelerators, and the serving system also needs memory for runtime overhead and the growing KV cache, so weights-only sizing is a floor.

A request becomes a sequence of tokens. A tokenizer maps text into numeric IDs, and prompt length matters because every input token consumes compute and contributes to the memory held for the active request. Every token then invokes the model weights: accelerators execute layers of matrix operations, often communicating across devices, and memory bandwidth, interconnect speed, batching, and software determine how much useful output the hardware produces.

The KV cache keeps context available. Previously processed attention state stays in accelerator memory so the model does not recompute the entire conversation for every next token; longer context and more concurrent users require more memory. Inference then repeats one next-token decision at a time: the model produces a probability distribution, selects a token, appends it to the context, and runs again. Tokens per second and utilization turn rented capacity into product economics.

What's inside: 6 components, their suppliers and sources #

Accelerator + HBM Parallel compute

The accelerator performs most AI math while on-package HBM holds model state and activations close to the compute cores. Cloud meters may expose a whole GPU or a partition.

What to watch. The accelerator name alone is insufficient: count, memory capacity, partitioning, interconnect, and software support determine what workloads the meter can run.

Host CPU General-purpose compute

Host cores run the operating system, prepare data, and coordinate accelerators. Inadequate host resources can leave expensive GPUs waiting.

What to watch. Compare CPU-to-GPU ratios across VM shapes and ask whether host bottlenecks or tenancy choices reduce accelerator utilization.

System RAM Host memory

System memory holds host-side data and software state. It is distinct from the HBM packaged with the accelerator and usually has a different capacity, bandwidth, and supplier exposure.

What to watch. Separate ordinary server DRAM from accelerator HBM; both matter, but HBM tends to carry higher bandwidth, tighter qualification, and different economics.

Local storage Data cache

Local NVMe can stage datasets and temporary state close to the accelerator. Persistent disks, shared filesystems, snapshots, and egress are often billed separately.

What to watch. A low VM headline price may omit the storage and data movement required by the workload. Compare the complete bill, not the compute meter alone.

Network interface Data movement

The VM network connects accelerators across servers and moves data to storage and users. High-performance training shapes may include specialized fabric that ordinary GPU VMs do not.

What to watch. Does the VM expose the scale-up or scale-out fabric required for distributed training, and are network or egress charges material to cost per workload?

Virtual machine meter Commercial bundle

The customer rents a named shape in a region under an offer type. The catalog rate is an observable retail meter, not a realized average price or proof of available capacity.

What to watch. Normalize only like-for-like meters and preserve whether a price covers a full VM, an accelerator add-on, or a multi-year reservation total.

Market participants7 companies #

Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.

Company Role in this layer Revenue Capex PP&E, net
GOOGL Google Cloud GPU and TPU rentals $119.8B $80.6B $321.2B
AMZN AWS accelerated instances $200.6B $173.0B $446.0B
CRWV Specialized GPU cloud $2.6B $14.1B $46.7B
DOCN Developer cloud and GPU instances $281.2M $81.6M $1.0B
MSFT Azure GPU virtual machines $331.8B $115.9B $313.1B
NBIS AI cloud and GPU infrastructure $529.8M $4.1B $5.6B
ORCL OCI bare metal and GPU clusters $67.4B $55.7B $100.0B

07Workload

A benchmark asks how quickly the system can finish a defined job. Training and inference depend on the model, data, accelerators, network, storage, framework, precision, and software. Measuring time or throughput shows what that complete system can deliver under a fixed set of rules.

What a workload costs

The same infrastructure can train a model, tune its behavior, serve responses, or generate video. The training and video scenarios begin with a hypothetical capacity reservation. The inference scenario begins with traffic and uses published system throughput to derive the required fleet. These are user-adjustable estimates, not invoices or estimates for a named company.

You are a frontier lab.

Train the next frontier model.

A large corpus is tokenized and streamed through a distributed training system. Every accelerator repeatedly updates the model weights until the run reaches its target, or a failure forces part of the work to repeat.

Adjust inputs

Calculated results

216.0M

$972.0M

This is a user-adjustable retail-equivalent compute model, not a reported lab budget. Frontier labs may own infrastructure or negotiate materially different economics.

[benchmark API β†—](https://intelligence.mts.now/api/compute/benchmarks). The inference proxy uses

[NVIDIA's published Llama 3.1 405B throughput β†—](https://developer.nvidia.com/blog/supercharging-llama-3-1-across-nvidia-platforms/)and

[AWS p5e Capacity Block pricing β†—](https://aws.amazon.com/ec2/capacityblocks/pricing/).

What's inside: 3 components, their suppliers and sources #

Model, data + quality target Benchmark definition

A performance number is only meaningful when the model, dataset, target quality, precision, rules, and benchmark release are fixed. Different workloads cannot be collapsed into one universal speed score.

What to watch. Does the comparison hold the workload and target constant, or is a faster headline actually measuring an easier model, lower quality, or different rules?

System under test Hardware + software

The measured system includes accelerator generation and count, nodes, fabric, storage, framework, precision, and software tuning. Scaling to more GPUs only helps when the rest of the system keeps up.

What to watch. How much faster does the workload become as system scale rises, and how much of that gain comes from better performance per accelerator instead of a larger hardware count?

Measured workload output Useful performance

Training benchmarks report time to a defined quality target; inference benchmarks can report throughput and latency. These results show what the complete system delivers under fixed rules.

What to watch. Translate performance into economics: system-hours, energy, and utilization required to achieve the result. Faster is valuable only if the incremental hardware and power cost are justified.

Market participants9 companies #

Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.

Company Role in this layer Revenue Capex PP&E, net
GOOGL Frontier workloads and custom TPU systems $119.8B $80.6B $321.2B
AMZN AI cloud, Trainium, and model demand $200.6B $173.0B $446.0B
AMD Accelerator and software alternative $11.5B $1.2B $3.4B
ANET Distributed-training networking $3.0B $84.2M $312.5M
AVGO Fabric and custom accelerator silicon $22.2B $481.0M $2.8B
CRWV Workload-optimized GPU cloud $2.6B $14.1B $46.7B
META Large-scale model training and inference $60.8B $49.1B $225.7B
MSFT AI platform, cloud, and model demand $331.8B $115.9B $313.1B
NVDA Accelerator and software platform $81.6B $1.8B $12.4B

08Economics

Companies pay for the infrastructure before it earns revenue. Purchase commitments and construction spending first appear as cash outlays, then as property and equipment, and later as depreciation. The return depends on when capacity enters service, how fully it is used, what customers will pay, and how much work the system completes.

Standardized SEC facts for 53 public companies show the timing of the buildout. Capital expenditures run far ahead of depreciation, so much of today's investment will reach earnings only over future years.

From capital to tokens

A GPU-hour tells you what the hardware costs to rent. It does not tell you how much work the hardware completed. That depends on throughput, utilization, latency, uptime, and energy efficiency. Capital buys energized capacity; software and operations determine how much of that capacity is productive; productive capacity generates tokens.

Megawatts measure a rate of power, so a tokens-per-megawatt comparison needs a stated time interval. The model below reports both tokens per second per megawatt and tokens per megawatt-hour. Any comparison still has to hold the model, workload, precision, latency, and quality target constant.

build the system facilities + compute

β†’

energize MW power available to equipment

β†’

run useful work uptime + scheduling + software

β†’

deliver tokens at a stated model and service level

Adjust inputs

Calculated results

1.2M tokens/s

43.2M tokens/MWh

37.8T

$5.28

Assumptions

What's inside: 4 components, their suppliers and sources #

Commitments + capital expenditures Cash investment

Purchase commitments and construction spending begin before assets produce revenue. Cash capex can lead delivery and placed-in-service dates by multiple reporting periods.

What to watch. How much spend is contracted versus discretionary, when will equipment arrive, and what portion of current cash outflow is still nonproductive construction in progress?

Property + equipment Productive asset base

Completed infrastructure moves onto the balance sheet as property and equipment when placed in service. PP&E includes more than AI compute and cannot be treated as a pure GPU inventory.

What to watch. What portion of asset growth is AI infrastructure, when does construction become productive, and how does asset turnover evolve as capacity ramps?

Depreciation + amortization Income-statement cost

Capitalized infrastructure reaches the income statement over its estimated useful life. Actual operating lives, utilization, and residual value determine the economics beyond the accounting schedule.

What to watch. How do reported useful lives compare with the period over which accelerator fleets remain productive, and how much future depreciation is embedded in the growing asset base?

Revenue, utilization + margin Economic output

The asset base earns a return only when workloads consume capacity at a price above depreciation, power, networking, support, and other operating costs.

What to watch. Does demand ramp fast enough to absorb new capacity, and are price/performance gains creating more revenue and margin than depreciation and operating expense consume?

Market participants9 companies #

Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.

Company Role in this layer Revenue Capex PP&E, net
GOOGL Cloud and TPU infrastructure economics $119.8B $80.6B $321.2B
AMZN AWS capex and accelerated-compute revenue $200.6B $173.0B $446.0B
CRWV GPU-cloud utilization and financing $2.6B $14.1B $46.7B
DOCN Cloud utilization and developer demand $281.2M $81.6M $1.0B
EQIX Colocation, interconnection, and leased capacity $2.6B $1.3B $25.2B
META Owned infrastructure and advertising returns $60.8B $49.1B $225.7B
MSFT Cloud capex and AI monetization $331.8B $115.9B $313.1B
NBIS AI-cloud capex and utilization $529.8M $4.1B $5.6B
ORCL OCI capacity, leases, and contracted demand $67.4B $55.7B $100.0B

Engineering and program management

$3.6B

βˆ’$34.0M

βˆ’$76.0M

βˆ’$86.7M

-0.9%

-2.1%

-2.4%

$69.4M

$169.2M

$99.9M

$1.0B

$459.7M

$167.4M

$617.1M

TPU accelerators and Axion CPUs

$119.8B

$40.8B

$112.2B

34.0%

93.7%

$4.3B

$84.9B

$80.6B

$55.9B

$321.2B

$13.6B

$18.0B

$7.7B

Trainium accelerators and Graviton CPUs

$200.6B

$27.5B

$62.6B

13.7%

31.2%

βˆ’$11.6B

$161.4B

$173.0B

$78.2B

$446.0B

$75.2B

$96.3B

Instinct GPUs and EPYC host CPUs

$11.5B

$6.2B

$2.0B

$2.3B

53.8%

17.3%

19.9%

$4.1B

$5.3B

$1.2B

$5.1B

$3.4B

$521.0M

$784.0M

$12.2B

High-speed and power interconnects

$8.8B

$3.5B

$2.6B

$1.8B

40.5%

29.5%

20.2%

$2.0B

$2.7B

$647.1M

$4.7B

$2.9B

$827.2M

$565.7M

AI Ethernet systems

$3.0B

$1.9B

$1.4B

$1.2B

62.9%

45.4%

40.0%

$2.7B

$2.8B

$84.2M

$2.3B

$312.5M

$46.7M

PCIe and CXL connectivity

$392.4M

$287.6M

$89.2M

$153.1M

73.3%

22.7%

39.0%

$134.2M

$162.3M

$28.1M

$111.5M

$119.3M

$7.6M

$44.4M

$181.7M

On-site fuel-cell generation $746.4M

$225.5M

$72.2M

$73.7M

30.2%

9.7%

9.9%

$47.4M

$73.6M

$26.2M

$2.5B

$401.1M

$13.3M

$129.1M

$0

Custom accelerators for hyperscalers

$22.2B

$15.4B

$10.8B

$9.3B

69.5%

48.6%

42.0%

$18.3B

$18.8B

$481.0M

$19.6B

$2.8B

$313.0M

$1.3B

Chillers and HVAC systems

$5.3B

$259.0M

$238.0M

4.8%

4.5%

βˆ’$15.0M

$79.0M

$94.0M

$1.4B

$3.1B

$1.3B

$560.0M

On-site generation and backup power

$20.5B

$4.3B

$3.6B

20.9%

17.5%

$4.9B

$6.2B

$1.3B

$6.7B

$15.6B

$1.2B

$728.0M

Networking and compute systems manufacturing

$4.7B

$577.5M

$458.3M

$368.8M

12.3%

9.8%

7.8%

$273.9M

$767.2M

$493.3M

$535.7M

$1.0B

$83.6M

$32.5M

Networking, optics, and systems

$15.8B

$10.1B

$4.0B

$3.4B

63.6%

25.0%

21.3%

$7.8B

$8.8B

$1.0B

$7.1B

$2.6B

$700.0M

$1.7B

Optical components and transceivers

$7.1B

$805.0M

11.3%

βˆ’$1.0B

$79.5M

$1.1B

$1.2B

$3.0B

$521.9M

$316.2M

Mechanical and electrical construction

$3.3B

$844.2M

$558.0M

$441.6M

25.9%

17.1%

13.5%

$1.2B

$1.5B

$288.8M

$1.9B

$653.9M

$38.6M

$315.8M

Firm and nuclear generation

$5.8B

$580.0M

$513.0M

9.9%

8.8%

βˆ’$968.0M

$1.6B

$2.5B

$697.0M

$41.2B

$870.0M

$505.0M

GPU-cluster operator

$2.6B

βˆ’$49.0M

βˆ’$626.0M

-1.9%

-24.3%

βˆ’$10.5B

$3.7B

$14.1B

$5.5B

$46.7B

$2.5B

$16.3B

High-speed connectivity and DSPs

$1.3B

$908.3M

$445.0M

$472.3M

68.0%

33.3%

35.4%

$407.0M

$464.3M

$57.3M

$1.2B

$101.6M

$34.6M

$25.4M

Backup generators and distributed power

$9.5B

$2.5B

$1.3B

$968.0M

26.1%

13.5%

10.2%

$1.4B

$1.8B

$438.0M

$3.2B

$7.0B

$563.0M

$141.0M

$90.0M

Enterprise AI systems and racks

$43.8B

$7.8B

$3.7B

$3.4B

17.8%

8.3%

7.8%

$3.1B

$4.1B

$963.0M

$11.6B

$6.9B

$758.0M

$745.0M

Developer cloud and GPU instances

$281.2M

$154.7M

$29.4M

$35.4M

55.0%

10.4%

12.6%

$75.3M

$156.9M

$81.6M

$767.0M

$1.0B

$96.6M

$479.1M

$41.0M

Regulated utility and grid delivery

$9.0B

$2.7B

$1.6B

30.3%

17.2%

βˆ’$2.6B

$1.5B

$4.1B

$2.1B

$130.0B

$1.9B

$1.3B

Switchgear and power distribution

$8.5B

$821.0M

9.6%

$1.2B

$1.6B

$446.0M

$483.0M

$4.7B

$671.0M

$789.0M

Electrical and mechanical construction

$5.2B

$1.0B

$547.3M

$403.7M

19.8%

10.6%

7.8%

$230.1M

$289.9M

$59.9M

$924.4M

$278.6M

$37.8M

$105.7M

Colocation, interconnection, and leased capacity

$2.6B

$665.0M

$479.0M

25.3%

18.2%

$1.8B

$1.3B

$979.0M

$25.2B

$1.1B

$1.4B

$8.2B

Optical and electronics manufacturing

$4.6B

$556.5M

$462.9M

$473.0M

12.0%

10.0%

10.2%

$4.2M

$256.7M

$252.5M

$346.7M

$615.1M

$68.4M

$4.0M

$132.5M

AI systems and liquid-cooled racks

$10.7B

$747.0M

$624.0M

7.0%

5.8%

$1.4B

$2.6B

$1.2B

$5.3B

$5.6B

$1.7B

$1.7B

Data-center engineering and delivery

$4.1B

$810.7M

$286.7M

$136.6M

19.9%

7.0%

3.3%

$291.0M

$352.8M

$61.7M

$1.2B

$311.6M

$67.6M

$114.7M

Cooling and building controls

$6.1B

$2.3B

$613.0M

36.8%

10.0%

$148.0M

$698.0M

$2.1B

$333.0M

$1.3B

Optical components for data-center links

$3.0B

$1.3B

$524.8M

βˆ’$6.9B

41.7%

17.4%

-230.1%

$300.1M

$751.4M

$451.3M

$2.0B

$1.2B

$128.8M

$33.8M

Custom accelerators and interconnect silicon

$2.4B

$1.3B

$339.4M

$34.5M

52.1%

14.0%

1.4%

$483.1M

$638.8M

$155.7M

$3.8B

$972.5M

$221.7M

$54.5M

Large-scale model training and inference

$60.8B

$18.8B

$15.8B

30.9%

26.1%

$15.0B

$64.1B

$49.1B

$15.5B

$225.7B

$12.4B

$2.4B

High-bandwidth memory and server DRAM

$41.5B

$35.1B

$33.3B

$28.2B

84.6%

80.4%

68.1%

$26.1B

$45.7B

$19.6B

$25.0B

$56.4B

$6.9B

$737.0M

Azure GPU virtual machines

$331.8B

$225.5B

$155.2B

$133.7B

67.9%

46.8%

40.3%

$67.0B

$182.9B

$115.9B

$20.9B

$313.1B

$34.3B

$21.9B

Data-center cooling systems

$874.1M

$182.0M

$74.8M

$73.9M

20.8%

8.6%

8.5%

βˆ’$5.0M

$41.4M

$46.4M

$95.3M

$536.1M

$20.7M

$27.8M

$27.0M

AI cloud and GPU infrastructure

$529.8M

βˆ’$611.7M

$82.5M

-115.5%

15.6%

βˆ’$3.7B

$384.8M

$4.1B

$3.7B

$5.6B

$417.9M

$845.4M

Enterprise data and storage systems

$6.9B

$4.9B

$1.7B

$1.3B

70.7%

24.2%

18.4%

$1.9B

$2.1B

$198.0M

$2.1B

$592.0M

$179.0M

$246.0M

Enclosures and electrical protection

$1.5B

$558.0M

$300.7M

$215.9M

37.9%

20.4%

14.7%

$211.1M

$268.7M

$57.6M

$256.0M

$447.9M

$34.2M

$33.0M

GPUs, Grace and Vera CPUs, NVLink

$81.6B

$61.2B

$53.5B

$58.3B

74.9%

65.6%

71.5%

$48.6B

$50.3B

$1.8B

$13.2B

$12.4B

$997.0M

$4.3B

$45.8B

OCI bare metal and GPU clusters

$67.4B

$20.6B

$17.1B

30.6%

25.4%

βˆ’$23.7B

$32.0B

$55.7B

$31.3B

$100.0B

$7.6B

$30.2B

Electrical distribution and control equipment

$311.7M

$95.3M

$64.1M

$52.2M

30.6%

20.6%

16.7%

$184.7M

$195.0M

$10.4M

$633.6M

$118.6M

$6.5M

$2.5M

Training data and checkpoint storage

$1.1B

$723.3M

$19.9M

$24.1M

68.7%

1.9%

2.3%

$111.8M

$180.2M

$68.4M

$837.8M

$613.9M

$38.8M

$231.0M

Transmission and substation construction

$9.6B

$1.5B

$694.8M

$451.4M

16.2%

7.3%

4.7%

$1.0B

$1.5B

$451.0M

$506.4M

$3.7B

$230.4M

$125.4M

Server and rack integration

$10.2B

$1.0B

$625.9M

$483.4M

9.9%

6.1%

4.7%

βˆ’$7.7B

βˆ’$7.6B

$133.8M

$1.3B

$607.7M

$38.4M

$378.1M

$10.1B

Power and data connectivity

$5.2B

$1.8B

$981.0M

$748.0M

35.6%

19.0%

14.5%

$2.2B

$3.0B

$832.0M

$1.2B

$4.5B

$758.0M

$491.0M

Chillers and thermal management

$6.4B

$1.2B

$925.7M

19.3%

14.6%

$1.6B

$1.7B

$156.1M

$1.8B

$2.4B

$207.4M

$824.7M

UPS, power delivery, and cooling

$3.3B

$637.9M

$497.8M

19.5%

15.2%

$1.6B

$1.9B

$285.9M

$2.8B

$1.2B

$223.5M

$82.0M

Methodology and sourcesMethods, caveats, and source register #

Counts. Facilities are distinct rows in Epoch AI's current site table. Source rows also include dated construction observations plus chiller and cooling-tower reference records; they are shown separately so a source-line count is never presented as a facility count. H100-equivalents are a modeled scale estimate and remain distinct from physical chip inventory.

Company coverage. Companies are mapped to layers when they hold a disclosed product, service, operating, or demand role. The mapping is an expanding coverage map, not a market-share ranking or an investment recommendation. Financial facts are standardized from SEC XBRL company facts; where an issuer does not report a standardized concept, the gap is recorded rather than imputed.

Claims and specifications. Every value is reported, estimated, or derived, with an evidence tier from Tier 1 (regulatory and audited filings) to Tier 4 (tracked third-party estimates). Specifications come from official reference designs; supplier lists describe ecosystems unless a disclosed bill of materials confirms the vendor.

Prices and benchmarks. Price observations are official retail catalog meters with their effective dates; they are not realized average prices and do not prove available capacity. Benchmark results are disclosed MLPerf Training submissions and are comparable only within one workload and release.

Capital-to-tokens model. The starting inputs are illustrative assumptions, not a facility estimate or NVIDIA product claims. Results depend on model, precision, workload mix, latency and quality targets, software, uptime, and asset life. The model includes capital but excludes electricity, financing, labor, networking charges, and other operating costs. The energy ratio assumes continuous draw at the stated IT capacity; cooling and conversion losses are excluded. Tokens measure output volume, not usefulness or intelligence.

View source register14 source families #

Source family Publisher Lane Tier Cadence State
SEC EDGAR and XBRL U.S. Securities and Exchange Commission company economics Tier 1 Per filing live
Official investor relations Covered public companies company economics Tier 2 Per filing or earnings release live
AWS Price List Amazon Web Services rental pricing Tier 1 Per catalog change live
Cloud Billing Catalog Google Cloud rental pricing Tier 1 Per catalog change credential-gated
Azure Retail Prices Microsoft Azure rental pricing Tier 1 Per catalog change live
MLPerf Training MLCommons performance Tier 1 Per benchmark release live
EIA electricity data U.S. Energy Information Administration power permitting Tier 1 Hourly to monthly credential-gated
FERC EQR and eLibrary Federal Energy Regulatory Commission power permitting Tier 1 Quarterly and per docket live
EPA ECHO U.S. Environmental Protection Agency power permitting Tier 1 Daily to weekly live
Utility and local planning records Utilities, grid operators, states, counties, and cities power permitting Tier 1 Per docket, agenda, or permit jurisdictional
Public procurement awards USAspending.gov and awarding agencies hardware pricing Tier 1 Daily live
OEM configurations NVIDIA, Dell, HPE, Lenovo, Supermicro, and peers hardware pricing Tier 2 Per product release live
Supplier disclosures Hardware and component suppliers hardware pricing Tier 3 Per earnings release live
Distributor observations Authorized distributors and resellers hardware pricing Tier 4 Daily jurisdictional

Nothing on this page is an investment recommendation. Values marked estimated or derived carry their stated assumptions; source status matters as much as the headline number.

Access and citationData API, snapshot, and citation #

The record behind every figure is available programmatically. All endpoints return attributed JSON.

- [/api/compute/drop-snapshot](https://intelligence.mts.now/api/compute/drop-snapshot) : the full immutable capture
- `mts://compute/summary` : Model Context Protocol resource

MTS Intelligence (2026). The Physical Stack Behind AI: an attributed record of AI compute. MTS Atlas, snapshot 2026-08-25. https://intelligence.mts.now/compute

Facility, cluster, and construction data: Epoch AI (CC-BY). Benchmarks: MLCommons MLPerf Training. Chip roles and units: NVIDIA product documentation. Prices: official Microsoft Azure and Google Cloud catalogs. Financial facts: SEC EDGAR. Hardware reference designs: NVIDIA documentation.

── more in #ai-infrastructure 4 stories Β· sorted by recency
── more on @nvidia 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/the-physical-stack-b…] indexed:0 read:35min 2026-10-01 Β· β€”