Cerebras Overclocks WSE-3 Waferscale Engine To Boost Inference Oomph In “Nexus” CS-4 Cerebras Systems launched its CS-4 'Nexus' systems with an overclocked WSE-3 Turbo waferscale engine, doubling clock speed to 2.8 GHz from 1.4 GHz while keeping the same 900,000 cores and 44 GB of SRAM, to boost inference performance. The company plans to double throughput annually through 2029 using the same Nexus rack design. Cerebras Overclocks WSE-3 Waferscale Engine To Boost Inference Oomph In “Nexus” CS-4 Everybody has been expecting for Cerebras Systems, one of the second-generation of AI hardware startups that has been trying very hard to compete with Nvidia on AI compute for years, to launch its latest systems, dubbed the CS-4, sometime this summer. I pegged the CS-4 announcement at August 2026 when I was contemplating what Cerebras might do with its $1.1 billion funding round back in October 2025 https://www.nextplatform.com/compute/2025/10/01/what-is-cerebras-going-to-do-with-that-11-billion-in-new-funding/1650974 , which was before the company went public and when it was not even clear that Cerebras would go public in 2026. It is always nice to make a prediction and to be right. While Cerebras is launching a new system and a new rack design that wraps around it with the CS-4 systems, which are codenamed “Nexus” and which lays the foundation for several more generations of machines from Cerebras, what it is not launching is a new WSE-4 compute engine as you might expect and I certainly did last October when thinking about what was coming down the Cerebras roadmap pike. As it turns out, the CS-4 machines are getting an overclocked version of the current WSE-3 waferscale compute engine, with the exact same 900,000 cores and the exact same 44 GB of on-wafer SRAM, and made using the same 5 nanometer processes from Taiwan Semiconductor Manufacturing Co. Well, actually, I would guess that TSMC is using an enhanced version of the N5 family of chip etching, given that this process node is much more mature here in 2026 than it was when Cerebras first delivered the WSE-3 engines in March 2024. Which is good for Cerebras and its customers. And so is an overclocked WSE-3 Turbo, as the new compute engine is called, which has the ability to do twice the work as its predecessor. Which begs the question as to why Cerebras didn’t overclock the cores and SRAM on its wafers more than two years ago. My guess is that the power and cooling technology that drives this clock speed, which I think has doubled to 2.8 GHz from the 1.4 GHz used in the plain vanilla WSE-3 engine, was not yet there. In the case of the CS-4 system, the compute wafer is essentially the same, but with twice as much power pumped through it from the wafer packaging and a little more than twice as much cooling to drive the 2X clock speed increase and keep it from letting the magic blue smoke out of the wafer. Here are the feeds and speeds of the four generations of CS systems: Just for fun last year, I took a stab at what a future WSE-4 compute engine might look like being employed in these CS-4 systems. It looks like the CS-4 Plus or the CS-5 might get a WSE-4 engine that is actually different from the WSE-3 or WSE-3T. I stand by the basic premise I outlined last fall that the waferscale engines are SRAM capacity limited, which is why it was taking multiple CS-2s and CS-3s to run inference on frontier models. It is probably up to dozens of CS machines by now, and perhaps more. I added the WSE-4 Harder option this time around, cranking the clocks up to the 2.8 GHz I think the WSE-3 Turbo runs at and also doubling up the bandwidth on the I/O subsystem to better balance 1 million cores with 320 GB of SRAM – so much better than 320 GB of HBM memory so far away from compute – with an aggregate of 248 PB/sec of bandwidth. I have high expectations for Cerebras, clearly. And so do its customers and the company’s engineers, and its top brass who presented these charts that can loosely be called roadmaps. Here is the first one presented by company co-founder chief executive officer Andrew Feldman: And here is one presented by co-founder and chief technology officer Sean Lie: So the Nexus racks with their backpack compute node design will be used for three generations at least, and every year between now and 2029, Cerebras will promise to speed up the throughput of those systems by 2X. That is considerably faster than Moore’s law improvements, and I wonder if they mean peak performance or effective performance. The latter is more likely, like for instance, driving more effective performance by adding more SRAM capacity relative to compute to get those cores to do a lot more work than they can do now because they are, like GPUs, starving for capacity. I strongly suspect that the WSE-3 cores were not even close to able to saturate the bandwidth on the SRAM, and the same holds true as things got doubled-up with the WSE-3 Turbo. Eventually, Cerebras will have to go to 3D SRAM stacking to get at least the capacity in line with compute. Cutting the number of cores and leaving the SRAM capacity alone or increasing it on a wafer is not palatable from a marketing standpoint – it is basically admitting that you undercut the memory capacity to sell more devices and knew it. As is the case with Nvidia and AMD GPU accelerators. The Nexus Rackscale System, Exploded The new Nexus rack is what is really interesting and innovative about the CS-4 systems, and to the extent that Cerebras could get 2X the work out of a wafer compute engine by just cranking the clocks is fine. One time. From here on out, it will be hard to reduce clock speeds below this new ceiling. There are probably a bunch of reasons why the new Nexus node design split the compute wafer and its host CPU an Epyc CPU from AMD from the banks of power supplies and network interfaces, but the big reason is that now these three different parts of the system can be upgraded independently. The NIC modularity will also allow Cerebras to allow partners OpenAI and Amazon Web Services to attach their own choice of NICs on the WSE engines to link them to the GPUs and XPUs they use for the prefill part of inference as well as the training of GenAI models. Here is what the Nexus rack looks like: The Nexus rack has three power shelves in the front and three CS-4 compute backpacks, which plug into the power shelves, in the back. The backpack is a vertically oriented modular server, and the design of the machine has 60 percent more manufacturing automation according to Cerebras. The Nexus rack has three times the compute per rack as the CS-1 through CS-3 machines, which had one 16U chassis per rack with one WSE and its host inside. In theory you could put two in a rack with room for scale out networking and storage – if you could cool it. The important thing, according to Lie, is that the Nexus rack is optimized for large scale clusters and has 50 percent fewer components and can be deployed up to 3X faster than prior systems from Cerebras. Here is an image of the backpacks exploding out of the back of the Nexus rack: Note: They are not spring loaded and a server compute complex weighing what is probably on the order of 100 pounds is not going to shoot out of the rack at you. Lie said that the idea behind the backpacks is that Cerebras can take a Nexus rack and its power shelves and ship them to customers, and then ship the SWE backpacks when everything is all set up. For whatever reason, this apparently saves some time. It certainly gives Cerebras some flexibility between incoming supply chain components and outgoing shipments to customers. Here is an exploded view of the CS-4 backpack: In the back of this image is the power module, with the WSE-3 Turbo module that lays on top of it and then a cold plate with an array of screws holding the complex all together. You can see two network interface cards on the backpack, one on the top and one on the bottom. Based on the specs, which say “new higher speed wafer links,” I think the wafer I/O module has six Ethernet ports running at 200 Gb/sec, which is twice the bandwidth per port as the 100 Gb/sec I/O modules used with the prior CS-1 through CS-3 systems, which might have had only one Ethernet I/O fabric interface or could have had two all along. The spec sheets and presentations were never clear on this. If there is a 2X radix increase in the NICs, that should mean that Cerebras can scale up networking between CS-4 systems beyond the top limit of 2,048 nodes in a single, data parallel cluster. However, we do not think Cerebras will support more than 2,048 nodes no matter what the topology and radix of the NICs. We have tried to confirm many of these network details with Cerebras but have not heard back as yet. The updated networking hardware supports a programmable and low latency packet pipeline and also supports direct wafer-to-wafer links. It is not clear how configurable these are in terms of how many ports are used to interconnect wafers Scale out networking to build larger CS-4 clusters as well as front end networks to link to users and storage are done through a partnership between Ethernet switch maker Arista Networks. Select customers can get early access to CS-4 systems now, and general availability will be later in the third quarter of this year. I will drill down into performance claims for the CS-4 machines in a separate story.