The Physical Stack Behind AI A new attributed record of AI compute infrastructure catalogs 83 facilities, 482 GPU clusters and 2,935 price observations, finding that official H100 rental prices range from $4.57 to $18.37 per accelerator-hour across 155 catalog observations — a 4.0× spread within a single hardware generation. The record estimates the largest facility, Colossus 2, at 946 MW, and reports that the four largest infrastructure spenders logged $425.2B in combined capital expenditures in their latest fiscal years against $130.7B of depreciation. NVIDIA has proposed bringing 800-volt direct current closer to the rack, a design the record says would cut conversion stages and waste less energy, though most operating data centers do not use that architecture today. The Physical Stack Behind AI An attributed record of where AI compute sits and what it costs: facilities, GPU clusters, power, the chips themselves, cloud prices, measured training performance, and the companies building each layer. What the record shows - The largest facility in the record, Colossus 2, is estimated at 946 MW . All twelve of the largest facility figures are estimates Figure 2 . - Dated construction records show how quickly the biggest campuses can grow, in some cases reaching hundreds of megawatts of IT power within roughly a year, with one announced plan extending to 1,925 MW by 2028 Figure 1 . - Official prices vary widely even for the same accelerator. H100 rates run from $4.57 to $18.37 per accelerator-hour across 155 catalog observations, a 4.0× spread within one hardware generation Figure 4 . - The four largest infrastructure spenders reported $425.2B of combined capital expenditures in their latest fiscal years, compared with $130.7B of depreciation. Much of today's spending will reach the income statement over the useful lives of those assets Figure 5 . The underlying record contains 83 facilities, 482 GPU clusters, 2,935 price observations, and SEC facts for 53 public companies. Every value is reported, estimated, or derived, with its source attached. How the system fits together Jensen Huang calls the AI economy a five-layer cake https://blogs.nvidia.com/blog/ai-5-layer-cake/ : energy, chips, infrastructure, models, and applications. NVIDIA calls the physical system within that infrastructure layer an AI factory https://blogs.nvidia.com/blog/ai-factories-the-new-infrastructure-of-intelligence/ : a data center designed to turn data and electricity into intelligence. A site takes shape in the order land → shell → power → cooling → compute. The sections below follow that build, then move inside the building from chips to racks and clusters, before looking at virtual machines, workloads, and economics. Build the site An AI campus is built in sequence: land, shell, power, cooling, then compute. We begin with the grid because the available power often sets the pace and ultimate size of the project. 01Grid + power A data center starts with a power connection. Generation, transmission, substations, and backup systems determine how much electricity the campus can use. A headline power figure might describe capacity that is contracted, permitted, energized, or already available to servers. Each marks a different stage of the project. NVIDIA has proposed bringing 800-volt direct current closer to the rack. The design would require fewer conversion stages, waste less energy, and leave more room for computing equipment. It describes a possible transition from today's AC facilities through hybrid systems to native 800 VDC sites. Most operating data centers do not use this architecture today. NVIDIA 800 VDC architecture ↗ https://www.nvidia.com/en-eu/data-center/technologies/800-vdc-architecture/ What's inside: 4 components, their suppliers and sources Grid supply Generation The generation fleet and wholesale market that supply energy to the local utility or balancing authority. Nameplate generation is not the same as firm capacity available to a continuously loaded AI campus. What to watch. Is the power physically deliverable and firm through peak conditions, or is the announcement only an energy purchase or aspirational generation build? Firming + backup Reliability Batteries, on-site generation, utility reserves, and backup systems bridge outages and variable supply. Backup capacity may be permitted for emergencies without being authorized for continuous operation. What to watch. How many hours can the system carry the IT load, what emissions or operating limits apply, and does the backup design support uptime without becoming stranded capex? Interconnection + substation Delivery Transmission upgrades, transformers, switchgear, and the interconnection agreement convert a grid promise into power that can reach the campus. This is often the schedule-critical path. What to watch. Is the project in a queue, under study, contracted, under construction, or energized, and who pays for network upgrades if scope or timing changes? Delivered campus power Usable capacity Gross utility service is reduced by power conversion, cooling, and other facility loads before it becomes IT power available to servers. The conversion ratio is usually summarized by power usage effectiveness. What to watch. How much contracted capacity is energized, how much reaches IT equipment, and how quickly can racks consume it without waiting for cooling or network completion? Market participants9 companies Hover or tap a company name for role, standardized SEC facts, derived margins, and sources. | Company | Role in this layer | Revenue | Capex | PP&E, net | |---|---|---|---|---| | BE | On-site fuel-cell generation | $746.4M | $26.2M | $401.1M | | CAT | On-site generation and backup power | $20.5B | $1.3B | $15.6B | | CEG | Firm and nuclear generation | $5.8B | $2.5B | $41.2B | | CMI | Backup generators and distributed power | $9.5B | $438.0M | $7.0B | | DUK | Regulated utility and grid delivery | $9.0B | $4.1B | $130.0B | | ETN | Switchgear and power distribution | $8.5B | $446.0M | $4.7B | | POWL | Electrical distribution and control equipment | $311.7M | $10.4M | $118.6M | | PWR | Transmission and substation construction | $9.6B | $451.0M | $3.7B | | VRT | UPS, power delivery, and cooling | $3.3B | $285.9M | $1.2B | 02Facility A data-center campus can contain several data halls, the large secured rooms where rows of racks are installed. The halls open in phases as their power, cooling, security, and network systems are commissioned. Almost every watt consumed by computing equipment becomes heat, so the cooling system limits how many racks each hall can support and how long they can run at full power. The twelve largest facility records, each unit square 25 MW of estimated facility power. All twelve are estimates, and several describe campuses still under construction; the largest, Colossus 2 at 946 MW, is roughly the electrical draw of a mid-sized city. What's inside: 4 components, their suppliers and sources Substation + electrical yard Power conversion High-voltage service, transformers, switchgear, UPS systems, and distribution equipment step grid power down and route it safely to data halls. What to watch. Transformer and switchgear lead times can gate energization even after utility capacity is awarded. Track redundancy and the difference between ordered and installed equipment. Cooling plant Thermal infrastructure Chillers, cooling towers, heat exchangers, pumps, and liquid distribution remove heat from dense racks. Rack architecture determines how much heat must be rejected to air versus liquid. What to watch. Does the cooling design support the rack density being purchased, and are water, heat-rejection, and mechanical permits aligned with the compute delivery schedule? Data halls Deployable floor Secured white space, busways, cooling distribution, and network pathways where racks are installed in phases. A campus announcement can include future halls that are not yet commissioned. What to watch. How many halls are shell-complete, powered, commissioned, and occupied, and which reported capex belongs to the current phase versus the full master plan? Control + network rooms Operations Building-management systems, security, telemetry, carrier rooms, and operations tooling keep power, cooling, and networks observable and available. What to watch. Physical completion is not the same as operational readiness. Commissioning, carrier diversity, monitoring, and trained operations staff determine when revenue-producing workloads can begin. Market participants12 companies Hover or tap a company name for role, standardized SEC facts, derived margins, and sources. | Company | Role in this layer | Revenue | Capex | PP&E, net | |---|---|---|---|---| | ACM | Engineering and program management | $3.6B | $99.9M | $459.7M | | CARR | Chillers and HVAC systems | $5.3B | $94.0M | $3.1B | | FIX | Mechanical and electrical construction | $3.3B | $288.8M | $653.9M | | ETN | Electrical distribution and protection | $8.5B | $446.0M | $4.7B | | EME | Electrical and mechanical construction | $5.2B | $59.9M | $278.6M | | J | Data-center engineering and delivery | $4.1B | $61.7M | $311.6M | | JCI | Cooling and building controls | $6.1B | $148.0M | $2.1B | | MOD | Data-center cooling systems | $874.1M | $46.4M | $536.1M | | NVT | Enclosures and electrical protection | $1.5B | $57.6M | $447.9M | | PWR | Electrical infrastructure construction | $9.6B | $451.0M | $3.7B | | TT | Chillers and thermal management | $6.4B | $156.1M | $2.4B | | VRT | Power and thermal infrastructure | $3.3B | $285.9M | $1.2B | Assemble the compute system Compute is assembled from the inside out. Chips sit in rack-scale systems with memory, networking, power, and cooling. Multiple racks are then linked into a cluster. 03Chips Inside each rack, GPUs and CPUs do different jobs. GPUs handle the model's parallel math, and inference throughput is usually measured in tokens. CPUs handle sequential work such as tool calls, code execution, and browser sessions. An agent can move between the two hundreds of times before it finishes a task. What the chips measure GPU throughput is usually measured in tokens. For agentic systems https://en.wikipedia.org/wiki/Agentic AI , CPU performance is easier to understand in completed tasks and concurrent agents. The denominator then tells you whether the claim is about energy, cost, serving demand, or the output of an entire site. The output of a GPU is tokens https://en.wikipedia.org/wiki/Large language model Tokenization , produced by parallel math https://en.wikipedia.org/wiki/Parallel computing on thousands of cores. The output of a CPU is tasks: the sequential steps around the model, such as tool calls, code, and browsers. A token count only becomes a claim once it has a denominator. Tokens per watt is a question about energy: how much model output remains after power conversion, cooling, and the chip's own efficiency are taken into account, which is the figure the power layer sets 01 Grid + power layer-power . Tokens per dollar is the commercial question: tokens divided by the meter you actually pay, whether a VM-hour, an accelerator-hour, or the depreciation https://en.wikipedia.org/wiki/Depreciation on hardware you own, and the same chip gives different answers under different meters 06 Virtual machine layer-vm . Tokens per user is throughput per person served: how many concurrent users one system holds at an acceptable speed, where batching, the KV cache, and latency targets all trade against each other the inference process inference-process . Tokens per AI factory treats the whole plant as one machine: NVIDIA's unit for a gigawatt-scale campus designed and operated as a single system, with its DSX blueprint as the reference design for one 02 Facility layer-facility . The CPU's units are tasks per second and agents per CPU. Token throughput describes the model running on the GPU. Task throughput describes the sequential work surrounding it. Public benchmarks for complete agent workflows are only beginning to emerge, so this page does not assign the CPU a performance figure. What's inside: 4 components, their suppliers and sources GPU + HBM Parallel compute Thousands of cores execute the same operation on different data at once, with high-bandwidth memory stacked beside the die so the math never waits on the wires. Its output is tokens. What to watch. A tokens-per claim needs its denominator stated: per watt is a power question, per dollar is a meter question, per user is a serving question, per AI factory is a site question. Host and agent CPU Sequential compute Fewer, faster cores for work that has to happen in order: tool calls, code execution, browsers, sandboxes, data pipelines, and orchestration beyond the model. NVIDIA's Vera CPU is purpose-built for this agentic work, with 88 custom Olympus cores and LPDDR5X memory. Its output is tasks. What to watch. Ask what percentage of an agentic workload's time runs on the CPU, and whether the host CPU is sized so the GPU is not idling at full price while a tool call finishes. The agent loop GPU-to-CPU link The model runs on the GPU, hands a step to the CPU, and the CPU runs it and goes back to the GPU to ask what's next. Agents work on a branchy decision tree, so one task can cross this link hundreds of times. A slow CPU step leaves the expensive GPU waiting. What to watch. Compare GPU-to-CPU ratios across systems and VM shapes; the ratio was designed before agentic workloads arrived, and agent-workflow benchmarks that would validate it do not exist yet. Units of output Measurement NVIDIA measures GPU output in tokens per watt, per dollar, per user, and per AI factory, where an AI factory is a whole gigawatt-scale site designed as one machine. CPU output is tasks per second and agents supported per CPU. What to watch. Two claims with the same numerator can be answering different questions. Restate every throughput number with its denominator before comparing it to anything on this page. Market participants7 companies Hover or tap a company name for role, standardized SEC facts, derived margins, and sources. | Company | Role in this layer | Revenue | Capex | PP&E, net | |---|---|---|---|---| | GOOGL | TPU accelerators and Axion CPUs | $119.8B | $80.6B | $321.2B | | AMZN | Trainium accelerators and Graviton CPUs | $200.6B | $173.0B | $446.0B | | AMD | Instinct GPUs and EPYC host CPUs | $11.5B | $1.2B | $3.4B | | AVGO | Custom accelerators for hyperscalers | $22.2B | $481.0M | $2.8B | | MRVL | Custom accelerators and interconnect silicon | $2.4B | $155.7M | $972.5M | | MU | High-bandwidth memory and server DRAM | $41.5B | $19.6B | $56.4B | | NVDA | GPUs, Grace and Vera CPUs, NVLink | $81.6B | $1.8B | $12.4B | 04Rack A modern AI rack contains far more than GPUs. The reference system shown here combines CPUs, GPUs and their HBM, networking, local storage, management hardware, power shelves, busbars, and liquid cooling. Every part affects how the rack performs, what it costs, and how quickly it can be installed. The rack is increasingly the unit of competition. CPU, GPU, HBM, networking, storage, power, management, and liquid cooling all affect the performance of the system. NVIDIA's DSX https://www.nvidia.com/en-gb/data-center/products/dsx/ blueprint extends that co-design to the whole AI factory, connecting reference systems, simulation, operations software, facilities guidance, and partner technologies. NVIDIA defines and validates the reference architecture, while its partners build and operate the sites. What's inside: 4 components, their suppliers and sources NVLink switch system Scale-up fabric Nine 1RU switch trays connect the 72-GPU NVLink domain through a passive copper cable backplane. Each tray contains two NVSwitches with 72 NVLink ports plus its own provisioning, telemetry, security, and control hardware. What to watch. Scale-up networking is part of the rack bill of materials and system yield. It is not captured by multiplying a standalone GPU price by 72. Compute trays Compute + host system Each of the 18 liquid-cooled 1RU compute trays contains two Grace CPUs and four Blackwell GPUs, plus cluster networking, BlueField DPUs, local NVMe, management controllers, and the operating-system image. What to watch. The deployable unit carries substantially more content than accelerators alone: CPUs, DPUs, NICs, storage, boards, cold plates, management, assembly, and support all affect price and lead time. HBM3e GPU memory High-bandwidth memory HBM sits in the GPU package and supplies model weights and activations at far higher bandwidth than ordinary server memory. NVIDIA specifies the aggregate rack capacity and bandwidth; it does not identify the memory supplier in this product specification. What to watch. Track HBM capacity per accelerator, stack generation, bandwidth, packaging yield, and supplier qualification. HBM availability can constrain GPU shipments and shift value toward memory suppliers. Power + liquid cooling Rack infrastructure Eight power shelves convert AC input to nominal 50–51V DC and distribute it over a busbar with N+N redundancy. Liquid manifolds and cold plates cool CPUs and GPUs while other components remain air cooled. What to watch. A roughly 120kW rack changes the facility bill of materials. Delivery is constrained by electrical distribution, liquid loops, commissioning, leak detection, and the facility’s ability to accept dense racks. Market participants9 companies Hover or tap a company name for role, standardized SEC facts, derived margins, and sources. | Company | Role in this layer | Revenue | Capex | PP&E, net | |---|---|---|---|---| | AMD | Accelerators and rack-scale systems | $11.5B | $1.2B | $3.4B | | APH | High-speed and power interconnects | $8.8B | $647.1M | $2.9B | | DELL | Enterprise AI systems and racks | $43.8B | $963.0M | $6.9B | | HPE | AI systems and liquid-cooled racks | $10.7B | $1.2B | $5.6B | | MU | High-bandwidth memory | $41.5B | $19.6B | $56.4B | | NVDA | Reference racks, GPUs, CPUs, and NVLink | $81.6B | $1.8B | $12.4B | | SMCI | Server and rack integration | $10.2B | $133.8M | $607.7M | | TEL | Power and data connectivity | $5.2B | $832.0M | $4.5B | | VRT | Rack power and liquid cooling | $3.3B | $285.9M | $1.2B | 05Cluster A cluster links many racks into one computing system. High-speed networking keeps the racks in sync, shared storage feeds them data, and scheduling software assigns the work. The cluster is operational only when all of those pieces have been commissioned and workloads can run reliably. Zooming back out from one machine to all of them: every cluster record with a first-operational date and a scale estimate, 2016 to present. The march up the log axis is the buildout. What's inside: 4 components, their suppliers and sources Scale-out fabric Inter-rack networking Ethernet or InfiniBand switches, adapters, optics, and cables connect racks into a training system. Fabric topology determines how efficiently additional accelerators contribute to a distributed job. What to watch. Does networking capex and optical content rise faster than accelerator count, and does the delivered topology provide enough non-blocking bandwidth for the target workloads? Compute racks Installed capacity Configured accelerator racks supply the physical compute. Physical chip count, rack count, commissioned capacity, and H100-equivalent capacity are different measurements. What to watch. Distinguish ordered, delivered, installed, networked, and operational racks. Usable capacity depends on the system around the chips and the power state of the facility. Shared storage Data plane Parallel filesystems and object storage feed training data, absorb checkpoints, and recover jobs. Slow checkpoint or input pipelines can leave the accelerator fleet idle. What to watch. Can the storage layer sustain workload throughput and failure recovery at cluster scale, and is storage or data movement becoming a material share of cost per useful token? Control plane Orchestration Schedulers, health monitoring, provisioning, and failure recovery turn a hardware fleet into a service that model teams can use continuously. What to watch. How much installed capacity is actually available and productively scheduled, and what software or reliability advantage lets one operator earn more from the same hardware? Market participants14 companies Hover or tap a company name for role, standardized SEC facts, derived margins, and sources. | Company | Role in this layer | Revenue | Capex | PP&E, net | |---|---|---|---|---| | ANET | AI Ethernet systems | $3.0B | $84.2M | $312.5M | | ALAB | PCIe and CXL connectivity | $392.4M | $28.1M | $119.3M | | AVGO | Ethernet switch silicon and custom accelerators | $22.2B | $481.0M | $2.8B | | CLS | Networking and compute systems manufacturing | $4.7B | $493.3M | $1.0B | | CSCO | Networking, optics, and systems | $15.8B | $1.0B | $2.6B | | COHR | Optical components and transceivers | $7.1B | $1.1B | $3.0B | | CRWV | GPU-cluster operator | $2.6B | $14.1B | $46.7B | | CRDO | High-speed connectivity and DSPs | $1.3B | $57.3M | $101.6M | | FN | Optical and electronics manufacturing | $4.6B | $252.5M | $615.1M | | LITE | Optical components for data-center links | $3.0B | $451.3M | $1.2B | | MRVL | Interconnect, optics, and custom silicon | $2.4B | $155.7M | $972.5M | | NTAP | Enterprise data and storage systems | $6.9B | $198.0M | $592.0M | | NVDA | Accelerators and scale-up fabric | $81.6B | $1.8B | $12.4B | | PSTG | Training data and checkpoint storage | $1.1B | $68.4M | $613.9M | Put capacity to work Once the cluster is commissioned, customers can rent the capacity and operators can measure what it produces. Workloads and utilization then determine whether all that spending produces a return. 06Virtual machine Cloud customers rent this hardware through virtual machines. Cloud providers expose this hardware as named instance types and charge by time. The price usually covers GPUs, host CPUs, system memory, local storage, networking, and platform services, so two hourly rates may include very different amounts of hardware. What cloud capacity costs: 2,935 price observations from official cloud catalogs, restricted to on-demand meters with a disclosed accelerator count so rates compare per accelerator-hour. Each hardware generation enters the catalog at a higher rate, and the same accelerator spans a wide range across regions and VM shapes, a 4.0× spread for the H100. A catalog rate is a public list price, not the average price customers actually pay. What happens when a model serves a request The model's weights must be loaded into accelerator memory before inference can begin. A request then passes through the same sequence for every token it generates. This explains what the machine is doing; the workload section below starts with traffic and estimates the fleet required to serve it. The trained parameters are the model. They begin as files on storage, and before inference can run those values must fit in accelerator memory at the chosen numerical precision