SpaceX logo (public domain) via Wikimedia Commons
Months of operating without backup power and cooling systems forced a strategic overhaul at facilities serving Google and other major AI clients.
SpaceX is overhauling how it builds and tests xAI’s data centers after its facilities in Tennessee and Mississippi ran for months without complete backup power and cooling systems. The result was exactly what you’d expect: outages, disrupted AI training runs, and uptime that fell below the 99.9% target the company needs to keep its biggest clients happy.
Those clients include Google, which signed a deal worth $920 million per month for access to roughly 110,000 Nvidia GPUs. When your monthly invoice looks like a small country’s GDP, you tend to notice when the lights flicker.
What went wrong #
The core issue was speed over stability. SpaceX had been racing to scale compute capacity, relying on makeshift power solutions like gas turbines and Tesla Megapacks instead of waiting for permanent backup infrastructure to be installed.
A power outage at the Memphis facility disrupted AI training processes, which is particularly painful given how expensive and time-consuming those workloads are. The outages reportedly impacted major clients including Anthropic and Google, both of which depend on consistent uptime to run their own AI operations.
SpaceX has since responded with corrective measures led by Elon Musk and his team. The new approach prioritizes installing backup power and cooling systems earlier in the construction process and running more thorough pre-operational testing before facilities go live.
Leadership shake-up and engineering reinforcements #
The reliability problems triggered personnel changes across the data center operation. Multiple data center executives departed, though the specifics of those exits remain unclear.
To fill the gap, SpaceX reassigned over 300 engineers from other SpaceX divisions to the data center effort after more than 1,300 employees expressed interest in transferring.
The scale of what’s being built #
As of June 2026, SpaceX reported 1.4 GW of compute capacity across its facilities. The target is to exceed 2 GW by year-end, which would represent roughly a 43% increase in about six months.
Some expansion plans have already been postponed to address supply chain challenges. Projected external compute rental revenue hit $2.6 billion for Q2 2026. The Google contract alone, at $920 million per month with a delivery deadline of September 30, 2026, creates an enormous incentive to get things right, as failing to deliver approximately 110,000 Nvidia GPUs on schedule would jeopardize what may be one of the largest recurring infrastructure deals in the AI industry.
The next six months will be telling. SpaceX needs to grow from 1.4 GW to over 2 GW while simultaneously retrofitting existing facilities with proper redundancy, meet the Google GPU delivery deadline, and do all of this with a partially new leadership team and hundreds of recently transferred engineers still getting up to speed on data center operations.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our