TL;DR — Key Takeaways
- Cerebras’ CS-4 combines three WSE-3T wafer-scale processors to deliver up to 750 petaFLOPs of AI compute.
- The system is designed primarily for large-scale AI inference, with Cerebras claiming speeds up to 30 times faster than competing GPU-based systems.
- CS-4 offers 129.6 PB/sec of memory bandwidth and wafer-to-wafer latency as low as 2 microseconds.
Cerebras Systems has introduced the CS-4 AI system, the fourth generation of its artificial intelligence rack-scale machine the company says can deliver inference speeds as much as 30 times faster than competing GPU-based systems.
The CS-4 is built around three next-generation Wafer Scale Engine processors and represents the first system based on Cerebras’ new Nexus rack-scale architecture. The company is positioning the system primarily as infrastructure for serving large AI models rather than as a conventional GPU alternative for general-purpose computing.
Cerebras says CS-4 can deliver 750 petaFLOPs of AI compute, 7.2 Tbps of I/O bandwidth and 129.6 PB/sec of memory bandwidth. The company also claims the system provides up to 10 times the throughput per watt of its previous-generation CS-3. The CS-3 also had just 125 petaFLOPs of AI compute power and about 21 PB/sec of memory bandwidth.
The heart of the system is Cerebras’s wafer-scale approach, an AI processor the size of a small pizza box. The Cerebras design, called the Wafer Scale Engine, is a giant processor rather than many small processors linked together by high-speed networks the way NVIDIA, AMD, and other accelerators function.
The WSE-3 processor measures 8.46 inches by 8.46 inches and has 900,000 cores and four trillion transistors. The typical GPU is about two inches square with billions of transistors.
NVIDIA and other accelerator vendors scale performance by linking multiple GPUs through high-speed interconnects, but communication between separate accelerators introduces additional latency.
For CS-4, Cerebras says wafer-to-wafer communication latency has been reduced to as little as 2 microseconds. That allows the system to distribute extremely large models across multiple wafer-scale processors without introducing the same level of latency as conventional accelerator clusters. The CS-4 also represents an attempt to make Cerebras’ unusual hardware easier to deploy. Cerebras says the new system has cut the number of components involved in deployment roughly in half. That could be important as the company attempts to move from relatively specialized AI installations toward large-scale deployments.
Cerebras recently went public and reported $180.1 million in revenue for the second quarter, up 74.3% from a year earlier, although the result fell short of Wall Street expectations. Cerebras subsequently raised its 2026 revenue forecast to between $880 million and $890 million.