How Meta built a 1GW data center in Ohio and is planning for a 5GW data center in Louisiana Meta built and operates Prometheus, a one-gigawatt-scale AI training supercluster in New Albany, Ohio, spanning thousands of acres with dozens of buildings and hundreds of thousands of GPUs, and now plans a five-gigawatt "Manhattan-sized" data center called Hyperion in Richland Parish, Louisiana, Meta Head of Infrastructure Santosh Janardhan said at the AI Infra Summit. Meta used data-center-scale tents of roughly two to two and a half football fields each, providing 30 megawatts and rated for level-one hurricanes, cutting deployment time by 70% to 80%, and reduced interruptions by a factor of 50 to target no more than two interruptions per 8,000 GPUs per day. Prometheus, which supports Meta's Superintelligence Labs effort to train next-generation large language models, remains under construction while serving training jobs and adding GPUs. Become a member of GB MAX to gain exclusive access to the industry and to the most influential global B2B leadership community in the business of gaming, entertainment, and tech. Join now https://go.gamesbeat.com/gb-max/ and also get a VIP ticket to GamesBeat Next Nov 2-3, SF . Meta recently described just how gargantuan its AI data centers are becoming. Even I was shocked at how big these places will be. Data centers are a vital part of the AI revolution, but they’re also raising a lot of concerns in local communities in the U.S. that don’t want them. Santosh Janardhan, Meta’s Head of Infrastructure, talked at the recent AI Infra Summit https://gamesbeat.com/computing-pioneer-david-patterson-talks-about-ai-and-the-future-chips-ian-cutress/ about how Meta built and operated Prometheus, a huge one-gigawatt-scale AI supercluster, or high-end data center in New Albany, Ohio. His aim was to take the audience behind the scenes. Now Janardhan said Meta plans to take the lessons of Prometheus to build a five-gigawatt scale AI supercluster in Louisiana. Prometheus serves as the foundation for Meta’s Superintelligence Labs effort to train and serve next-generation large language AI models. Meta built and operated Prometheus, a distributed AI training cluster spanning thousands of acres. The campus includes dozens of buildings and hundreds of thousands of graphics processing units GPUs . And Prometheus is far smaller than Hyperion, which Janardhan said will be a “Manhattan-sized” data center in Richland Parish, Louisiana. Prometheus challenged the assumption that a training cluster must fit under one roof. In this case, the buildings were separated by three to five kilometers, with latency governed by distance at roughly five microseconds per kilometer, Janardhan said. Rapid construction Meta deployed data-center-scale tents measuring roughly two to two and a half football fields, each providing 30 megawatts and built to withstand level-one hurricanes. The tents reduced deployment time by 70% to 80%, shifting construction from years to months, while requirements were narrowed from a conventional data center to those needed specifically by a training cluster. I have to wonder how these tents will deal with tornados that are common in the Midwest. At hundreds of thousands or millions of components, failures are statistically continuous. Early Prometheus deployments suffered interruptions from changing network, cooling, hardware, and construction variables, including concrete dust affecting fiber links. Meta reduced interruptions by a factor of 50, targeting no more than two interruptions per 8,000 GPUs per day. PAF, or parallelism-aware fault tolerance, uses model parallelism, data layout, physical topology, availability, cost, and latency to route around failures and recover from continuous checkpoints. The next data center Prometheus remains under construction while serving training jobs and adding GPUs. Lessons from this multivariate production experiment are informing Hyperion, planned as a five-gigawatt, Manhattan-sized data center in Richland Parish, Louisiana. The central open question is how effectively these clusters are used; GPU count, fiber distance, and facility size are ultimately vanity metrics unless they improve user and business outcomes, Janardhan said. “I lead the infrastructure team at Meta. It’s a fancy way to say that if your Facebook, your Instagram, your WhatsApp is not working, I’m the person” that is responsible, Janardhan said. “Not in long ago — I want to say maybe a couple of years ago in 2024, we were creating our first AI training cluster, and it was decent size.” It was about 100,000 several GPUs based on Nvidia’s H100 technology. It was pretty beefy for 2024. “When we were in the process of building it, we were creating our clusters, the Llama-sized clusters on it. It was clear that the scaling laws were going to hold pretty well, so we had to go and create our next big cluster, which was going to be 10 times that size,” Janardhan said. So Meta started building out the next one. Overall, Meta’s infrastructure includes multiple gigawatts and tens of millions of servers. “We run a substantial portion of the internet, but AI infra is different. It is not conventional infra that you see everywhere,” he said. “You have to think and rethink a lot of the things that you have done for the last 10, 20 years, right? So what we started building out is something we call Prometheus.” To call it a data center is doing it a disservice, he said. “It’s probably the most ambitious construction networking project that we have ever attempted. And to put things in perspective, we were also using the cluster when we were building it,” he said. “Now we’re talking about thousands of acres. We’re talking about 30 to 40-plus buildings. We’re talking about hundreds of thousands of GPUs. It’s a fairly big-sized sort of apparatus, as you can imagine.” Making assumptions and then changing them To build something of this magnitude, you have to make a lot of assumptions, and in this case, some of them held up and some didn’t, he said. “If you think about good innovation, good engineering usually happens when you hit roadblocks. I think that’s what we ended up doing,” Janardhan said. The first assumption was that an AI cluster lives only under on roof. The GPUs should all be next to each other to be able to communicate. But with Prometheus, the campus started offas a cluster of buildings, but then the ambition grew. “So this creates a massive networking problem. If you run a AI training cluster, especially if we do it at north of 100,000 GPUs, what happens is that sheer throughput of the cluster, the amount of bits you can push through it, is as fast as the slowest connection you have. And when the connections are different, you kind of are stuck. Physics doesn’t change much. Distance equals latency,” Janardhan said. “If you think about the speed of light, it’s about five microseconds per kilometer. If you take a mile, it’s maybe eight microseconds. You do a round trip, 16, right? So that that is not something you can change.” The second thing about AI training is that it assumes no data loss. It does not deal well with data loss at all. You need to create a lossless network, and this is not how the internet is built,” Janardhan said. “The real internet actually assumes sort of data loss, so you have to deal with these two constraints. And then you have to build out something along the distances that I was mentioning earlier.” Meta did this by creating its own new network layer. Meta called it the backend aggregation network. This was like an “Ethernet super spine.” “There’s a spine that connects all the data hauls within a building. There’s a spine that connects all the buildings, and then it essentially aggregates all the subspines in a three to five kilometer radius. “And when you go back to back, you have really thick network pipes. That’s the best way to describe it. Now, this gives you two important advantages. The first one, at the back layer, it gives you headroom. That means that you can go and deal with data loss. You can deal with packet loss at the local back layer, which is a huge deal because because then you can concentrate it locally and fix it locally,” he said. The second thing is that it is not one spine. It is multiple spines. But he said, “You do not want a a single point of failure SPOF in your network, and this gives you multiple points where you can do end to end and actually make more reliability in the system.” The back-to-back network wiring runs in the tens of petabytes a second. To put that in perspective, you can probably transmit the Library of Congress in less than 30 seconds along these pipes. “It’s just humongous pipes, so it makes it truly seamless. The people using the network wouldn’t have any clue where the clusters are situated,” Janardhan said. There are two fabrics inside the network itself. One is the DSF. Think about as as a fabric, he said. It disaggregates the network traffic, chops it up into packets, assigns it links, reaggregates it at the destination. The other fabric is the NSF, which is just cheap network switches. “We take a lot of them. Spray the packets. Gives us great control at the NIC level. We can route traffic. We can mold traffic, and together NSF and DSF actually create what we call the back ,” Janardhan said. If you look at the wiring that connects he buildings, there are a series of pipes, and each of those pipes has tens of thousands of fibers. That’s as much as four million kilometers of wires. “You can go to the moon and back four times. The scale is just off the charts,” he said. Janardhan noted that data centers are big, honking constructions that take time to build. “Whenever Mark came to me and said, ‘Hey, I need a data center,’ I said, ‘Well, you should have talked to me two years ago,'” Janardhan said. So the team started studying how they could build faster. One idea was to put the servers into tents. “Now, this is not your garden variety tent. I just want to be clear about this. This is not the one you go camping in. These tents are the size of two and a half football fields. They are built to the standards of a data center, and they can withstand hurricanes of a level one, and so they are resilient,” Janardhan said. Just by putting the servers in tents, the process becomes 70% to 80% faster. “You can build them up in the order of months, not years. Faster time to deployment, faster time to cluster, faster time to model, faster time to profit,” he said. “And we didn’t just build these structures here. We also fundamentally went and re-examined what was into a data center, and what does what does it actually need for the training cluster.” The team stripped down the requirements from a data center to what a cluster requires, and that cut down the time as well. The team managed to deploy what it needed. Each of the tents can handle 30 megawatts worth of computing. “You can throw the biggest internet party on the planet in one of the tents,” he said. The third assumption A third assumption was that once you build a data center, it’s done and the work is over. So what was the third assumption? Well, one of the things that we talked about was hey, once “Well, I wish life was so simple, right? Training clusters are fundamentally a gigantic, complex, distributed systems problem. That’s what it is.” Each component within that cluster is subject to failure. If you have a component that fails once every 10 years, that’s resilient. But if you take 10,000 components and put them together, then one is failing every so many weeks. The resultant cluster is not reliable. So you have to design the system for failure. So the network has to be modified as needed and that means the data center changes. You have to constantly improve the cooling and improve other things at the same time. That adds up to even more complexity. “Clusters are sensitive. What happens if you have a single interruption in a cluster, a network interruption or any other error ? The training job can stop, or at the very least slow down. Now, if it slows down, the throughput is not happening,” he said. He added, “We worked with the vendors, and I’m pretty proud to say that we brought this down by a factor of 50, and our North Star here is about two interruptions per 8,000 GPUs per day. We could not afford more than two interruptions per 8,000 chips. That is an extremely high level of reliability that you want to aspire to,” he said. The team is operating within that target now. He added, “Now, little things compound. We were building so fast. We were putting the servers down so fast. We were moving servers before the concrete on the servers on the on the floor hopefully cured. We had concrete dust issues and fiber optic links. Little things you didn’t think about this. That means you had random interruptions within our training job.” Concrete dust became something the team studied and had to deal with. “The thing to understand is at this scale, failures are going to happen. It’s important to understand the patterns. It’s understand important to understand you need to model out the failures, the issues that can happen, and then have a plan to tackle it because you will have problems,” he said. Software was the easy part? The fourth assumption was that software is the easy part. “Now this is a mea culpa. I’m one of the software guys, and we should have done better, right? Now at the end of the day, if you think about software, I’ve been talking about a lot of complexity here. Talking about network, training, distances. New products are being introduced. The job of the software thing is to bundle up all of that and present it as an urgent layer that I can go and work on and operate,” he said. He added, “If I am a researcher in the lab, I do not care how pretty your data center is. I do not care how complex it is. I have a window, I have a prompt. I’m hitting enter. The damn thing is not running. The job of software is to obfuscate the complexity and to make the thing just work, Our job is to actually obfuscate a bit of the complexity.” He said the infrastructure team is appreciated when things work, and people notice you only when things don’t work. “When you turn on that faucet, you want the water to come out. When you flip the switch, you want electricity to come out. When it doesn’t work, people notice. When it works, we will take it for granted,” he said. We want the Matrix to just flow.” The software’s job on the cluster is to be aware of the physical layout of the cluster and optimize it based on availability, cost, latency — and it can route around failures. If you hit a failure, life will go on. “Think back to the time when all of you were in school, right, and the science teacher probably taught you how to do a science experiment. How would you do a science experiment? You take an experiment, you change one variable at a time, run the experiment, and that makes it easy to diagnose. It makes it easy to figure out what went wrong, what went right, and then you go and change the next variable. We do not have that luxury in real life. I wish, right?” he said. “Life would be so much better. What we end up doing is running this massively multivariate experiment in production. That’s essentially what Prometheus is,” he said. Now his team will take the lessons to the next cluster. That next big cluster will be Hyperion, the five gigawatt cluster at a single location in Parish, Louisiana. “We think it’s the biggest data center ever constructed on the planet. This thing blows things out of the water. It’s the size of Manhattan. It is absolutely humongous. It takes hours just to drive around it, right?” he said. “That’s a story for another time.”