# Presentation: From Fab To Token - The State Of The Market

> Source: <https://www.infoq.com/presentations/ai-hardware-tokenomics/?utm_campaign=infoq_content&utm_source=infoq&utm_medium=feed&utm_term=global>
> Published: 2026-08-18 16:00:00+00:00

## Transcript

**Jordan Nanos:** My name is Jordan. I work at SemiAnalysis. I'm on the technical staff. I primarily work on a project called ClusterMAX, which I'm going to talk about towards the end. The topic of the presentation is From Fab to Token. I'm going to try to take you guys through a little bit of what SemiAnalysis does in terms of AI research and semiconductor supply chain research. The purpose of this is a little bit different than some of the other talks. I'm a practitioner. I use the models. I develop software. I do this in service of analysis of hardware. I would say I'm a hardware guy. We care a lot about testing the new GPUs, measuring performance, testing cloud providers. What I'm going to try and do here is give you a sense of how understanding the hardware can help inform developing great software and vice versa. To me, it's really hard for me to use AI and feel like I really understand what's going on without having any understanding of the chips, the data centers, the systems, even the entire supply chain that goes into it. I'm going to try and take you guys through some of that today.

## Who is SemiAnalysis?

SemiAnalysis is a semiconductor and AI research firm. We're the number one substack in the technology sector. Our three-tier newsletter goes to 280,000 subscribers. We have over 85 team members now. For what it's worth, I joined last summer. I was employee number 31, so we're growing really quickly. We generally attend 100-plus conferences a year like this one to try and learn and keep track of what's going on in the industry. The core thing that we do is sell research products. We sell subscriptions to like a specific feed. If you're interested in data centers or chips or tokenomics or wafer fab equipment or anything like that, both investors and industry companies will subscribe to this research. You get emails in your inbox and you get financial models and charts and things like that.

The key question that I'm going to ask and try to answer in this talk is what is the choke point? Currently, we think that there's a lot of people who are asking questions like this, maybe also asking the question of, is this a bubble? I'll try and walk through and answer that question a little bit. I'm going to cover four different things in the market right now. First is chips. What are the constraints? How are we producing them? Two is data centers. How are we getting power and data center facilities online to be able to put these chips? Third is system performance. This is the core of what I do at SemiAnalysis. I work on a project called ClusterMAX. We also have one called InferenceX. In ClusterMAX, we test hands-on all of the cloud providers in the industry. This is the hyperscalers, neoclouds. We test over 100 of them hands-on, write up our findings, give out that research for free.

Then on InferenceX, what we do is we test the chips themselves. We compare NVIDIA, AMD, pretty soon TPU, Trainium, many of the startups. We run open-source models that are popular and run on all the latest hardware. I also give that for free. That's at inferencex.com. You can check out that data right now if you're interested in those sorts of topics. Then the fourth thing is end user demand. We call this tokenomics, the economics of inference. We can talk a little bit about ChatGPT growth, Claude Code versus Codex, what's going on with OpenAI and Anthropic? Where's the value occurring across the stack, and things like that.

## Part 1: Chips

Part one is chips. I'm going to show a lot of charts in this presentation. First of all, hyperscalers are spending a lot of money on chips right now. If you look at this chart, what you can see is that the yellow bar is the consensus CapEx from analysts, and the blue is what they're actually going to spend in 2026. Google, Amazon, Meta, and Microsoft have revised up how much money they're spending mainly on chips, but also the CapEx that goes into data centers, an absolutely massive amount. Many people may have seen the figure of $1 trillion going forward. A lot of this money is obviously going to chips. No y-axis on this chart, but this is some of our institutional research to show you how much money is actually going towards AI accelerators a year. It's ramping really fast. The key takeaway here is really that TSMC is not keeping up with the demand growth.

If you look at how many wafers TSMC is producing every quarter, you could just look at this chart. It's not exponential the way the others are. The breakdown on the slide here is across the different process nodes. The yellow is the old stuff, 6 nanometer, 7 nanometer, 4 and 5 is in blue. Red's the 3 nanometer, which is the leading-edge node right now. You can also see 2, which is coming soon. The reason why is because TSMC is taking a much more measured approach to their capital expenditure than what we've seen from the hyperscalers. You can blame this on cultural issues or people who have been here before think it's another bubble, but this is effectively the key constraint, which is that you can't get more wafers, wafers being the thing that chips are built from, without more expenditure on fabs, without more expenditure on equipment coming from TSMC.

While they are increasing CapEx, this little bump here in 2022, where they came down towards the end of COVID, we're seeing in the market right now, which is that they can't meet the demand for what people are looking for. All of this growth at TSMC is because of AI. On this chart, you'll see that in the yellow, this is the non-AI shipments in terms of demand of 3 nanometer wafers. The way that all of this has worked for a long time is that non-AI, in other words, smartphones and PCs were the leading edge. You'd get phones with the latest process technology before you'd get chips, like data center chips. This has completely changed. The vast majority of all 3-nanometer capacity at TSMC is going to be going towards AI. This comes at the expense of smartphones. If you're interested in maybe a takeaway here, you want a new iPhone, or specifically, if you want some of the Chinese phones, probably buy them right now, because they're about to get really constrained next year.

Again, if you want to see the breakdown by accelerator vendor, NVIDIA is in yellow in terms of their demand for wafers from TSMC. Broadcom is in blue. Broadcom produces the chips for Google, the CPUs. Annapurna in black is the division of Amazon that produces the Trainium chips, black or navy blue. There's marginal stuff on top. A lot of people maybe see us posting online or see other discourse about NVIDIA versus AMD. You can just see what's the demand for AMD there. They're a really small sliver of purple in the top. Very small just in terms of NVIDIA. If you think that companies in the startup ecosystem or whatever are going to take share from NVIDIA, it really depends on how much they can get from TSMC. It just doesn't look like that's going to happen. NVIDIA is going to be the lion's share of TSMC 3 nanometer revenue going forward.

AI growth is coming at the expense of smartphones. This is a specific number that shows how wafers that were originally targeted towards smartphones have changed. 3 nanometer wafers in terms of how they're being reallocated towards the R200, the latest Rubin GPUs that NVIDIA is going to be shipping towards the end of this year, and the TPU v7 from Google. This is also impacting memory. Maybe people have seen this stuff online about SK Hynix employees buying Ferraris recently or things like that. If you look at progressively over time how much of the world's DRAM capacity was being purchased for AI, it starts as like the bulk of it is not, 12% is AI back in 2023. Going into next year, we're going to see that over 60% of total DRAM wafer capacity is going to go into AI systems. It's impacting much more than just HBM, which is the traditional part.

If you look at how much of the total wafer production at TSMC is going to have HBM on it, HBM being like a core component of the NVIDIA GPUs, AMD GPUs, TPUs, HBM being High Bandwidth Memory, the percentage of total memory wafers that are going to HBM is increasing significantly over time. Maybe from a geopolitical level, higher level, what happens when we see this? In other words, when TSMC increases their CapEx or there's other political pressure to bring chip manufacturing into China from their government or potentially chip manufacturing facilities for memory in Korea, we can look at these three countries and see from public data how much equipment they're actually importing. They all have import-export data that they publish. This is a product of ours called ChipBook. It's one of the entry-level products. The three major tools in wafer fab equipment, lithography, deposition, and etch.

I'll go back to litho. You can see on the total equipment imports, that yellow line popping up is China importing lithography machines, potentially ahead of regulations from the U.S. and other countries around the world. Also, Taiwan and Korea are having this inflection up where I made that comment earlier, TSMC brought down CapEx at the end of 2022, but now they're ramping it back up. People are rising to meet this demand. We expect to see this to continue to grow over time. This is roughly the same shaped graph, whether you're looking at litho, deposition, or etch equipment.

In terms of what it means for actual practitioners, people who use GPUs in the cloud or use tokens that are produced by GPUs, it just means that there's new GPUs coming. This is a screenshot from NVIDIA's GTC presentation. You can see a big lineup of Rubin GPUs that are going to be shipping towards the end of this year. This is all based on HBM4 and 3 nanometer capacity from TSMC. So far, what we've seen is that the performance improvements from the new GPUs have been immense. You can say that this is due to a whole bunch of different factors going into the system. I'll just summarize it by this. When we look at inference performance in InferenceX, which I'm going to talk about a little bit more later, what you can see is that there's a significant reduction in the cost to serve tokens over time.

There's also a significant improvement in the speed at which tokens are produced, when you compare like a Hopper generation system, like an H100, to a Blackwell generation system, like the GB300 NVL72. On the left side, you can see a tweet that we put out where we described this and said, at GTC in 2024, Jensen had these charts. We charted up to Jensen Math saying GB200 NVL72 is going to be 35 times faster. They generally are misleading, we would say lying. What happened, our founder, Dylan, literally sent Jensen an email after we got hands on with the GB200 systems towards the end of last year and kept doing some testing. They bragged about this at their GTC Conference on stage during the keynote this year, which is that we'd realized more than 35 times of a performance improvement, up to 50 times. I can show you a chart.

This is a live chart that I took a screenshot of. They're on the right side of the screen. This y-axis is a cost per million tokens. This x-axis is the interactivity, meaning the token throughput per user. Left side is like slow 40 tokens per second. Right side goes up above 100 tokens per second. This is the fast mode. You can see that there's a sharp increase in the cost it takes to serve fast tokens on the H100 there. Going up to over $2 per million tokens using an 8K, 1K, like 8K token input, 1K token output workload shape. Meanwhile, the GB200s are down here at less than 10 cents per million tokens. That's clearly a factor of 20 times. This was incredibly surprising. This is like a cherry-picked benchmark. There are benchmarks that they can cherry pick that make it look like it's over 35 times faster.

The point is that if you look at the specs of the GPU and you see that flops increased by less than two times going from Hopper to Blackwell, memory bandwidth increased at a little over two times, memory capacity at two, two-and-a-half times, what was driving this? It's really the fact that a lot of frontier model inference is networking constraint. Networking is the performance bottleneck. That blue line that shows you going from aggregate scale-up bandwidth within the rack, just on the specs from 3,600 total to 64,000 means that any software technique that was constrained by the networking, specifically things like prefill-decode disaggregation, where you deploy multiple copies of the model, and then you do KV cache transfer between the copies. You can do pre-fill, processing the big long input context, and then decode, processing the output, creating the output on separate GPUs with separate configurations that are tuned for those workloads.

The only way you can really do this is with a high bandwidth network. Maybe some people have heard the description of co-design. As new hardware enables new software to be developed, we can realize these sorts of performance improvements. It's shocking the fact that we thought they're doing their classic lying thing. Actually, Jensen responded to Dylan's email. He didn't say this on stage, but Dylan's email said, "When we tested the performance and saw it wasn't just 35 times faster, but it was 50 times faster, we thought you were lying." Jensen responded with, "I never lie."

What's coming next in networking is really CPO, Co-Packaged Optics. This is a screenshot from Lambda, a neocloud that's installing the first version of these optical switches. These are going to be scale-out networking switches. A lot of this technology for optics is already used within the systems or in long range. I'll show you the tradeoffs basically. Today, for scale-out networking, you use copper. That means that you have these DACs, ACCs, like active copper wires or active electrical cables that you use to actually build these clusters. A really big trend in the industry right now is moving to optical networking in this space. You can see my nice AI-generated slides here explaining that stuff. Conceptually, from a developer's perspective, if previously you could use 72 GPUs in a rack or a world size of 64 of those 72 GPUs, the next generation systems from NVIDIA that are going to be shipping into next year, the first ones this year with Vera Rubin are still going to be 72 GPUs in a rack.

Next year, there's 144, 288, 576. There are these bigger domains that are multi-rack. The 576 one specifically needs these optical cables. Co-Packaged cables is this really interesting technology where you can potentially save a lot of power. You could be saving up to 20% of your total power on the cluster that you consume by using this technology. That also means in terms of raw input materials, it's lower cost. The thing that's a bit in question is, what's the failure rate going to be? In other words, people currently have an experience of paying a premium for higher reliability and they may be continuing to make that trade until this technology is more proven.

## Part 2: Data Centers

Moving from chips and then networking at the end and really the description of what the constraints are in the systems, now we're going to talk about data centers. The key question that comes to everybody's mind is like, let's assume that you can produce a whole bunch of chips, which definitely is happening. NVIDIA is going to have $500 billion of revenue this year and next year. Where do all those chips go? I think this is making its way into the political zeitgeist right now. You can just look at this chart and see that excluding China, in South Asia, there's still a whole bunch of data center capacity that's coming online. If there's one takeaway from this presentation, I would say that people should assume that going into the future, we're going to have a lot more tokens, a lot cheaper on higher quality models. There's nothing that leads us to believe that there's a significant constraint in terms of how much training or inference compute is going to be deployed in the U.S., or how the chips are going to perform or how the models are going to perform.

**Participant 1:** That's a very surprising result. I think all of the charts that led up to that moment might have indicated to me that we're going to see extremely high cost runs on tokens and compute and memory and everything. I'm very surprised to hear that you think there's not going to be a constraint on build-out in the United States. I was getting from the other charts that there likely would be, and I've been worried about how much of my 401k I should spend on NVIDIA appliances before token winter comes. If that's not going to be a problem, thank you.

**Jordan Nanos:** I think that maybe the key takeaway from this presentation will be that TSMC is the major constraint. We will be able to solve a lot of the other bottlenecks in terms of data centers, in terms of networking performance, in terms of end user demand that people may be nervous about when it comes to a bubble or threat of a bubble. Token winter is an interesting term. There's plenty of tokens going around right now with just the current deployments of GPUs. This section of the presentation will give you a sense for how big some of the data centers that are coming online are. I think you just have to forecast that in the future, which is like when hyperscalers are pouring in, they're taking their free cashflow to zero. They're deploying all these GPUs in all these data centers. That just means there's going to be a lot more tokens this year and probably next year, likely through to 2029 as well.

Look at this chart. Again, no y-axis, but this is our data center model. We track 6,000-plus data centers globally. We do satellite imagery to see where the construction progress is at. We track all the permits that they file for. We think we have a good understanding of what projects are real and what projects are not. What you can see is this red bar in terms of total surplus of self-built data center capacity is currently forecasted to increase over time, which just means that we're going to have more data centers empty waiting for chips by the end of 2028 than we will have chips to fill them up with. That's generally the high-level conclusion. That's not to say that this is like a done deal. There's a lot of delays going on. There's a lot of political changes across the world. When one project gets canceled in the U.S.

and another project gets ramped up in Malaysia, things just even out over time. I'll walk you through a little bit of why this is the case and why a lot of people who talk about the constraints on the construction of the facilities, constraints on the electrical equipment, or constraints on the power and cooling equipment, or constraints on the power equipment are not necessarily going to be correct about the bottleneck. This is a chart that explains it. Data center building going back to 2010 in non-AI segments, meaning like not GPUs, is the yellow bar at the bottom. This is continuing to grow. This is your general-purpose cloud workloads. You can see that data centers coming online are quite massive in terms of the total power consumption that they're going to have. They're in the blue at the top of this chart. You can also see exactly what they spend on.

This is a little preview from our data center model to say, if somebody is spending CapEx on a data center, what is it actually going towards? What's the bulk of the expense? You can see that a huge component of that up here is UPS generator and switch gear and power distribution. This is like the components that produce the power that the chips consume. That's like the orange-yellow-white bar at the bottom. Then some of the blue stuff is the cooling systems. This is your CRAC, your CRAH, and your CBUs, as well as your chillers, cooling towers, pipes, valves.

I went to one of these sites with a couple guys on the team. This was in Lake Mariner near Buffalo, a town called Barker, New York. They're building 750 megawatts of compute capacity over there, over time. It's a company called TeraWulf. A former crypto miner who took over a coal mining power plant. The thing with these sites is previously you exported the electricity. You just switched the way the power flows, and now you can bring in electricity, bring in power from the nuclear power plant that's nearby, as well as from the grid with NYSEG. On the left, it's us standing outside a 400,000 square foot facility. The Buffalo Bills are putting up a new stadium. This is the biggest construction project in New York State right now. On the right side, you can see me standing next to one of the buffer tanks. This is a small facility.

By their standards, it was 20 megawatts. They have four of these tanks of just water that's available in case something happens with the liquid system, and they can run it without actually needing to produce more chilled water. It's unbelievable the scale at which this is happening. Again, 750 megawatts. I'm standing next to something that's powering a 20-megawatt data center. I'll cover this in a bit. Companies like Meta, Amazon have already built 1 gigawatt facilities, and there are projects in flight right now in the U.S. and in India and other places around the world that are multiple gigawatts of scale. Maybe for what it's worth, San Francisco as a city consumes less than a gigawatt of power. New York's up at 9 gigawatts or something like that. Unbelievable the scale at which this is happening.

The thing about the new chips that makes these data center facilities a little crazy is the fact that the rack power has gone exponential. Back in the day, before SemiAnalysis, I worked at a company called HPE, Hewlett Packard Enterprise. We sold servers to many banks and telcos and federal government stuff, and they would just be typical servers. The typical rack at that time was 9 kilowatts per rack. Then in 2017, 2018, we saw V100 GPUs start to drive this a little bit higher to 12 or 14 kilowatts per rack. That's the low bar there. Fast forward to end of next year, we're going to be expecting racks that are 600 kilowatts each. I stood next to racks that are 190 kilowatts. This is 42U high, like a standard rack that's in every data center. It's consuming 20 times the power currently and going up to 40 or 60 times the power in the future.

That drives things like this. New GPUs are coming. Tracking Michael Dell's cover photo there. What's that rotated? You can see it's a Vera Rubin rack. This is the new GPUs that are going to be shipping in December. The big improvement from Blackwell. A lot of new technology going in there. This is what's taking the rack power consumption from 130 to up to 190 kilowatts or 200 kilowatts per rack. The thing about this that's impacting the industry is that you need a different facility level electrical system in order to power these racks. Instead of stepping down the power to like 415 or 480 volt, what we're going to see in these racks is that you have to deliver power to the rack at 800 volt in DC, not three-phase AC power. There's a huge transition that's happening right now in the data center industry as people try and work on first retrofitting, which is the yellow bar at the bottom, and then blue at the facility level.

By 2030, our current forecast is that almost 80% of all data centers globally are going to be using this new electrical system. A lot of innovation happening on the power side. The problem that people rightly point out is that the grid is sold out. If you want to bring a gigawatt online right now, you can't really get a load like interconnection agreement signed with many of the utilities. You can see on this chart, it's unbelievable. There are new large load requests that we track in our energy model in blue, and you can see which have been approved in orange. It's just not happening. People are not approving new data centers. The solution is to actually do your own power generation behind-the-meter. These are overhead satellite photos from xAI site called Colossus in Memphis, where you can see that they've brought on turbines on site to generate the power that they need to run the data center. If they can't get power from the grid, they just do it themselves with these turbines. All they really need is an air permit, maybe some noise permits to be able to burn all of this gas.

In general, a little teaser of what we do with the data center model. I also just want to comment on how fast some of these sites are going up. We track construction progress with satellite images and permits, and you can see with all of these pictures, there's a 2024 picture on the left and then a 2025 picture on the right. You can get a sense just how quickly some of this construction happens by looking at these progress photos over time. People who talk about labor constraints or machinery constraints or things like that are really not seeing how fast some of these sites are going up and how modular the designs of the buildings are now.

**Participant 2:** Just to clarify, are those sites in the United States?

**Jordan Nanos:** Yes, all four of these are in the United States here that I'm showing, but it's satellite, so we track stuff in Europe and Asia as well. Maybe the biggest problem for some of this stuff is actually just literally finding the site. Once we have the coordinates, we're able to direct satellite imagery and figure out how quickly construction's happening and how quickly these companies are disclosing delays, which is not always that quick.

A key question that a lot of people ask because of all this is like, are AI data centers actually increasing electricity bills for American households? I think the answer in general to cut to it is generally no. Anybody who lives in ERCOT, which is Texas's grid operator, and PJM, Pennsylvania-New Jersey-Maryland. PJM is where all the data center alley stuff is from U.S. East, and then Texas is where all the big stuff is going for Stargate and a bunch of the other projects. You can just see what consumer electricity prices are going towards here. We're not seeing an exponential ramp-up. I think that the big potential issue that people are forecasting is the inability to forecast demand. Again, the big reason for this is that so much of the data centers that are coming online, as I described earlier, are going behind-the-meter with their generation, which means they don't actually run on the grid.

They don't run or compete for the same power generation resources that are powering households right now. If you care about the bill breakdown, we break this down in an article about this, where we go into details. There are two aspects to this. One is that you could be drawing from the grid when there's no capacity available, in which case they need to spin up on-demand capacity to have a demand response. Then the second one is when there's a capacity charge. The breakdown is different because, effectively, when there is a capacity charge, you need to pay people to generate the stuff.

This is maybe one of the final slides to make the point about behind-the-meter generation, which is to say, how much are we forecasting? The biggest reason that consumer electricity prices could potentially increase is because the grid operator does not have a forecast for how much demand they're going to see in the next summer. They don't then sign the contracts with the power producers to be able to produce this demand. What we show here on this slide is the forecast that are coming from companies like PJM in blue compared to our forecasts in yellow. All you have to do is just track which data centers are coming online to be able to forecast what the total load's going to be, which we do pretty smoothly, pretty clearly, and are able to accurately predict this. Whereas PJM on their own grid are not able to predict this, and they under-build and then over-build in terms of how much power they have.

That's what spikes consumers' electricity prices. It's like, if they get that wrong and they need to scramble to find new generation to accommodate for the summer months or the winter months, that's when they need to pay a significant premium for people to produce power. Final slide on data centers, behind-the-meter is now the path. If you care, the typical interconnect queue that you need to wait for if you're bringing a data center online to use the grid is anywhere from two-and-a-half years to five-and-a-half years. Texas is the fastest. PJM is the slowest, roughly speaking, across the major stuff there.

## Data Center Security

**Participant 3:** What is the thought these days around the security of the data centers and the vulnerability as our workloads move more to AI and are reliant on these data centers? How well are they protected from drone attacks, missiles, those kinds of things, like something we've seen in the recent past?

**Jordan Nanos:** We talk to a lot of these data center operators who take the physical security and protection of their data centers quite seriously. This is both for theft of IP, as well as sabotage of the data center, as well as literal drone strikes and rocket attacks. This is more common in some places of the world than in others. It's a new area that people are exploring more of. Data centers today, I mean, the one that I was at, there's armed guards walking around, there's dogs. They take your license to get in. They do real security. It's a building, it's currently being built. It's not really running, maybe they're running some production workloads, but let's say on the security perspective, it's a 50-50 chance that they're actually running any of the frontier model weights in any of the data centers that I was next to. They're likely doing research for new stuff there.

I think it's a real thing that people consider. I think that the bulk of investment that I've seen right now has been on software security, as opposed to the physical security of the physical site. It's scary to see people target data centers as vulnerable places in a war. They're going after critical infrastructure and it's like the places where you produce the oil and then the data centers, places where you produce the tokens.

## Energy Sources for Behind-The-Meter Build-Out

**Participant 4:** Could you comment on the behind-the-meter build-out? What is the breakdown of the energy sources? Is it mostly natural gas and some nuclear? Is it other things? Can you just comment on how that breaks down?

**Jordan Nanos:** It's mostly natural gas. People have other creative ways of doing this. Generally, nuclear is not behind-the-meter if they're using nuclear today, but you may have seen public disclosures like Microsoft has signed a big agreement to restart Three Mile Island for some data centers for them. Others have plants and small modular reactors. Nuclear just takes a while to be produced, so it's not really a big component of it today. There are some sites that use hydro. There are some sites overseas that use coal, but most of it for behind-the-meter stuff is natural gas.

## Constraints on Turbine Production

**Participant 5:** There's a constraint on the key for number of turbines that [inaudible 00:38:08].

**Jordan Nanos:** There's constraints on how many of those turbines all the vendors can produce, but people are getting creative. One of our articles, we go through all those details on turbines, fuel cells, some of the companies that are converting jet engines into turbines, basically. Boom aerospace was going to create a fast plane, and now they're producing engines for data centers because they had a spec. I met a startup recently called American Turbines that says they're going to just produce turbines for data centers. It's like six guys in San Francisco with a warehouse. I don't know how they're going to do it, but good luck to them. Fuel cells are super interesting from Bloom Energy, not Boom Supersonic, two different companies, both producing power for data centers behind-the-meter. Fuel cells, a lot like no emissions, despite natural gas being the input still, or less emissions. This is where your tokens run.

It's like heavy industry. Actually, maybe one last point on data centers. When I say rack power density and stuff, it's absolutely fascinating how large, on a percentage basis, the amount of the data center itself is going to cooling equipment and electrical equipment, as opposed to literally the racks. As the chips get in the racks, get more and more dense, the white space, like the actual data hall in the data center today, maybe like 20% of the facility, like 80% of it is just these massive machines and pipes to produce the chilled water and the electricity you need to power these things. That's just going to increase over time. We're going to have these massive facilities that are football fields in size, and then you're going have two tiny data halls at the end of it. If you're wondering why people are talking about putting data centers in space, that's the motivation. That's not coming anytime soon.

## Part 3: System Performance

System performance. I'm going to talk about the constraints here in terms of how easily can you use these GPUs. The first project is ClusterMAX, where we compare neoclouds. This is a critical decision if you're training models or if you're running inference. Then, InferenceX, which compares the chips. Which chips should you be using if you're trying to run your business on top of these GPUs? First of all, what's a neocloud? It's getting a lot of play recently. People use the term. SemiAnalysis coined the term in October of 2024 in an article called, "AI Neocloud Playbook and Anatomy." At the time, there was a bunch of new companies that were converting their data centers that they previously used to mine crypto into hosting GPUs for people to train their models. These companies like CoreWeave or Crusoe or Lambda, they were growing like crazy. There were people using a lot.

The consideration was why. We tried to answer that question. If you are a company that's trying to get access to GPUs, particularly for a training cluster, you have a lot of companies that you can pick from. In our first version of ClusterMAX, which was in March of 2025, we did hands-on testing of 26 providers and rated them. In ClusterMAX 2.0, we did 84. Then 2.1 minor version update, we added 6 to get to 90. We're currently working on ClusterMAX 3.0, hopefully shipping by the end of this summer with a nice open-source repo I'll describe here. The point was to give people some perspective into the quality of these providers. Until you've used a cluster and worked with the support engineers from a company like CoreWeave or Oracle or Nebius or Fluidstack or Crusoe, you don't really have a perspective on why people might even pay a premium to run there when compared to hyperscalers like AWS or Google or Azure and many of the other companies. We track 200-plus of these companies now who are converting their facilities and installing GPUs. How do you sift through that noise? That's really what we try to answer with this research.

What makes a good neocloud? I think that there's basically two ways to look at this. The first is what your expectations are for the product you're going to be purchasing. For training, a lot of people still use Slurm. For inference, you always use Kubernetes. Sometimes you can use Kubernetes for training too. Then some people just look for standalone machines. We try to rate companies for all of those three scenarios with a set of criteria where we'll check that Slurm is configured properly, we'll check that Kubernetes is configured properly. Not to say you couldn't fix these things in a few days, a little headache, but if a company gives you something that works well out of the box, it's an insight into the fact that they know what they're doing. We also assess their monitoring dashboards and their health checks because such a key component of these clusters is reliability.

Some example tests that we run are included here. I'm not going to go through these, but I left these as takeaways for people to see what sort of testing we do, which results in an assessment. Of course, the ratings actually go beyond hands-on testing, which is to say we talked to over a hundred companies that are using GPUs. If people here are using Slurm clusters or Kubernetes clusters with any of these providers, we'd love to talk to you about your experience. We get insight into reliability and support experience at scale in a unique way there when people are willing to share what they're up to, as opposed to just our testing experience when we have the CTO on the phone or something. If you go to the website, https://www.clustermax.ai/tco, what you can do is actually plug in your own values for goodput and things like that to actually see what performance is like and what reliability is like at scale. You can say, this is the quote I'm getting from these people for the GPUs, for the storage, for the networking, and how many failures I expect to see on this lower tier provider compared to the better one where I expect to see less interruptions in my jobs and less failures on these GPU nodes.

InferenceX is the other project, I talked about a little bit already. InferenceX is an open-source benchmark that we say moves at the same speed as the AI software ecosystem, which is to say we run the latest nightlies of vLLM or SGLang or TensorRT when they get released. We work closely with those communities to integrate them into our repo on our hardware. We run across all sorts of different GPUs, and you really get to see what the performance is on a given GPU, on a given model. We cover open-source models like DeepSeek, Kimi, Qwen, MiniMax, things like that. I talked about how Jensen likes it. The whole repo for InferenceX is open source. You can see exactly what we're running, modify it, fork it, run these performance benchmarks against endpoints as you see fit. I won't cover this in detail. There's lots of content online where we talk about how to analyze these things. Beyond performance, we're also doing a lot of work on reliability, which is to say, when a new version of this software ships and claims a performance improvement, is it actually impacting the quality of the model? You have to run these evals against the model to make sure that things aren't changing there.

Final section, tokenomics, where does the value go? You can see that to us, tokenomics is the summation of a whole bunch of different research. Our TCO model where we assess how much the components that go into the servers actually cost. Our data center model where we figure out how much the data center costs to build and to operate. InferenceMAX where we see how much performance you're going to get on a given GPU, on a given model. Then this tokenomics model that I'm going to describe here where we actually make an assessment of end user demand and use some of our internal data to see this.

## Part 4: End User Demand

Claude Code has arrived. We track how many GitHub commits are being authored by Claude Code, how it tags itself in there. This was 4% of all GitHub commits back in January. It's only grown since then. Hundreds of thousands of commits every day from Claude Code. Incredibly fascinating. We wrote an article in February called, "Claude Code is the Inflection Point." I think it still holds up. It has driven an inflection where OpenAI now has less ARR than Anthropic. Anthropic is growing faster than OpenAI now. Amazon is a key beneficiary of this growth. We pointed this out in a recent article. You can see AWS in blue. Their operating margins are improving over time. Microsoft's operating margins are going negative. GCP's are staying flat. Why is Amazon improving their operating margin? Because Anthropic buys Bedrock. They're using a managed token as a service platform to power a lot of applications.

Many people are consuming through Bedrock where they have a revenue share agreement. What that means is that as Amazon improves the performance of a cloud model on their Trainium chips, they're just going to get more margin. They just continuously work on inference performance and they continuously improve the margins. Here's a chart that shows Bedrock is growing like crazy. This is our forecast to show Bedrock revenue as a percent of the total AI and total AWS revenue. As Anthropic grows, so does AWS. They're not quite so exposed to Gemini and Azure Foundry, even though the models are available on those services as well. Neoclouds in this space are largely irrelevant. They're a purple bar, you can't even see it, whereas the blue and the green are growing like crazy. If you're curious about where to use Anthropic's models and who's going to have capacity, we really expect this to continue because AWS there in blue is adding a bunch more data centers for our tracking, just in terms of like megawatts coming online.

Their projects are on time. They're going bigger than others. That's going to continue into next year. Just look at how fast they're building. These are Anthropic AWS AI training data centers in the bottom in 2024 and in the top in 2025. One year span. Look at how fast those buildings go up. These are 100-plus megawatt facilities in total. It's unbelievable how good these guys are at operations.

We believe agents are just getting started. We did some global economic TAM analysis to see across a bunch of different sectors what the pie that we think an agent can eat would be. It's massive. I want to clarify maybe two things as I wrap here. One is that we believe currently OpenAI and Anthropic are highly profitable. A lot of this stuff about subsidization may or may not be true depending on how much cash you're doing and how good they are at serving these models in terms of performance. We believe that they're very profitable just based on the disclosures that they've made during their funding announcements in terms of raising money. Anthropic hitting $47 billion of ARR recently and OpenAI's last disclosure being $2 billion a month, which puts them up at $24 billion ARR, even though they're much higher than that based on our tracking.

We think we're right. This is one of my peers, Joey, who works on the tokenomics team and shows that we're on point there. Then the final point is that we believe that the shift from training to inference is roughly a lie. As these companies raise more money, they're able to buy more training data centers and improve model performance. When people say like there's a big shift from training to inference, we really just don't see that in the data at this point.

**See more presentations with transcripts**
