# Cerebras details CS-4 rack and targets 10,000 tokens per second with CS-5

> Source: <https://runtimewire.com/article/cerebras-cs-4-nexus-rack-cs-5-token-speed>
> Published: 2026-08-26 00:30:51+00:00

# Cerebras details CS-4 rack and targets 10,000 tokens per second with CS-5

**Cerebras is using modular power, cooling and compute to chase an annual speed upgrade with its wafer-scale systems.**

By [Ryan Merket](/author/ryan-merket)
· Published

Primary source: [Cerebras on X](https://x.com/cerebras/status/2092384486302306384)

## Why it matters

Nexus makes Cerebras' rack, rather than a single wafer, the product roadmap. That could shorten upgrade cycles, but CS-5's 10,000-token target remains a 2027 promise.

[Andrew Feldman](https://datascience.uchicago.edu/people/andrew-feldman/?ref=runtimewire)'s [Cerebras](https://www.cerebras.ai/?ref=runtimewire) on August 25 detailed the rack architecture behind [CS-4](https://www.cerebras.ai/cs4?ref=runtimewire), a week after introducing the AI system, and previewed a CS-5 successor targeted for 2027. The update, shared during the Hot Chips conference, lays out Feldman's plan to turn wafer-scale processors into a repeatable datacenter platform rather than redesigning the surrounding machinery for every chip generation. ([Cerebras' CS-4 announcement](https://www.cerebras.ai/blog/introducing-cerebras-cs-4?ref=runtimewire))

That system-level focus follows Feldman through years of infrastructure work. Before founding Cerebras in 2015 with Gary Lauterbach, Michael James, Sean Lie and Jean-Philippe Fricker, Feldman co-founded energy-efficient server maker SeaMicro, which AMD acquired for $357 million in 2012. He had previously worked at Force10 Networks, acquired by Dell for $800 million, and Riverstone Networks. Feldman holds a BA and an MBA from Stanford University. ([Feldman's University of Chicago biography](https://datascience.uchicago.edu/people/andrew-feldman/?ref=runtimewire))

The throughline is efficiency gained by controlling the whole system. Cerebras built its first three generations around processors spanning most of a silicon wafer, avoiding the network of separate chips and switches used in GPU clusters. With CS-4, Feldman's team is applying the same co-design logic to the rack, including its power conversion, liquid cooling and I/O.

### A server rebuilt around the wafer

CS-4 combines three WSE-3 Turbo processors inside Nexus, a new rack platform with one processor in each rear-mounted compute backpack. Every backpack contains its own power conversion, cooling, I/O and control electronics. The shared power equipment sits at the front of the rack, allowing technicians to install or replace the compute assemblies from the rear.

Cerebras says the backpack has 50% fewer components than the previous design, uses 60% more automated manufacturing and can reduce deployment time from days to hours. Those deployment figures come from Cerebras rather than published field data, but the architecture addresses a practical constraint that benchmark charts usually skip: specialized accelerators still have to fit into datacenters built around standardized power, cooling and service procedures. ([Cerebras' CS-4 announcement](https://www.cerebras.ai/blog/introducing-cerebras-cs-4?ref=runtimewire))

The power-delivery design is particularly aggressive. According to Cerebras' [Hot Chips technical post](https://www.cerebras.ai/blog/ultrafast-frontier-inference-cerebras-deep-dive-at-hot-chips-2026?ref=runtimewire), CS-4 places AC/DC conversion roughly 0.5 millimeters from the wafer, compared with about 50 millimeters in conventional GPU systems. Cerebras says removing the printed circuit board from the final delivery path lets CS-4 send nearly twice as much power to the processor at almost the same voltage while limiting resistive loss.

Each CS-4 backpack can house up to 30 PSU modules and connects to the rack's liquid-cooling network, with leak detection and quick-disconnect valves. The layout separates high-voltage equipment at the front from water and compute at the rear. ([ServeTheHome's CS-4 rack analysis](https://www.servethehome.com/cerebras-talks-going-rack-scale-with-their-wses-at-hot-chips-2026/?ref=runtimewire))

CS-4 also supports disaggregated inference, where one type of processor handles the prompt-processing, or prefill, phase and Cerebras handles token generation, or decode. Cerebras has named AMD Helios and AWS Trainium as possible prefill platforms. The approach lets Cerebras sell its speed advantage without requiring customers to replace every part of an existing AI cluster.

### The speed claims need their denominators

Cerebras calls CS-4 the industry's fastest AI accelerator and says it can deliver up to 30 times the inference speed of GPU systems across a selected set of models. Cerebras attributes the comparisons to Artificial Analysis and its own August 2026 tests. Its published disclaimer says results can vary by model, workload, configuration and testing date.

Cerebras also claims more than 1,000 tokens per second on models exceeding 10 trillion parameters, wafer-to-wafer latency as low as two microseconds, up to twice the performance of CS-3 and as much as 10 times the throughput per watt. Some of those figures rely on internal benchmarks, projections or extrapolations rather than independent production comparisons. ([Cerebras' CS-4 announcement](https://www.cerebras.ai/blog/introducing-cerebras-cs-4?ref=runtimewire))

RuntimeWire [reported on August 13](/article/openai-gpt-5-6-sol-ultrafast-cerebras-preview) that Cerebras hardware was powering an OpenAI [GPT-5.6 Sol](/models/openai/gpt-5.6-sol) preview at up to 750 output tokens per second, with access limited while capacity expanded.

That speed matters most for coding tools, reasoning systems and agents that make a chain of model calls. Latency accumulates at each step, so faster decoding can shorten an entire task rather than merely making text appear more quickly on screen.

### CS-5 is a 2027 target, not a finished product

Cerebras says CS-5 will use a next-generation Wafer-Scale Engine in the same Nexus platform. The 2027 system targets up to 10,000 output tokens per second per user on models including [Gemma 4 31B](/models/google/gemma-4-31b-it:free) and [gpt-oss-120b](/models/openai/gpt-oss-120b). For multi-trillion-parameter models, Cerebras is targeting up to 5,000 tokens per second per user, 3 million tokens per second per megawatt and support for models exceeding 50 trillion parameters.

Those numbers are forward-looking targets. Cerebras has not presented them as shipping performance, and its technical post places them inside the legally required forward-looking-statements section. Nexus is the concrete part of the roadmap: Cerebras can retain the rack's power, cooling and I/O structure while changing the processor generation. ([Cerebras' Hot Chips technical post](https://www.cerebras.ai/blog/ultrafast-frontier-inference-cerebras-deep-dive-at-hot-chips-2026?ref=runtimewire))

Cerebras also previewed CS-6, a later design intended to combine wafer-scale SRAM and compute with stacked DRAM. Cerebras says work on that concept began in 2024. The planned memory expansion would let more of a large model reside within each system, reducing traffic between separate processors and racks.

### Feldman's architecture now has a public-market test

Cerebras completed a $6.4 billion gross IPO in May 2026, following a [$1 billion Series H](https://investors.cerebras.ai/news-releases/news-release-details/cerebras-systems-raises-1-billion-series-h?ref=runtimewire) at an approximately $23 billion post-money valuation. Tiger Global led that round, joined by Benchmark, Fidelity Management & Research, Atreides, Alpha Wave Global, Altimeter, AMD, Coatue and 1789 Capital. ([Cerebras' IPO announcement](https://investors.cerebras.ai/news-releases/news-release-details/cerebras-systems-announces-pricing-initial-public-offering?ref=runtimewire))

The capital gives Cerebras room to manufacture systems and build inference capacity. It also raises the standard for proving that wafer-scale performance can produce a broad, repeatable business. Cerebras reported $510 million in 2025 revenue, up 76% year over year, but its prospectus said MBZUAI and G42 represented 62% and 24% of that revenue, respectively. First-quarter 2026 revenue was $193.4 million. ([Cerebras' SEC prospectus](https://www.sec.gov/Archives/edgar/data/2021728/000162828026035214/cerebras-424b4.htm?ref=runtimewire))

CS-4 shipments are scheduled to begin in the third quarter of 2026. The meaningful test will come as customers run production models and comparable data emerges on speed, power use, reliability and cost. Nexus gives Feldman a credible mechanism for increasing performance without rebuilding the entire rack each year. CS-5 will show whether Cerebras can keep that cadence after the first modular system leaves the factory.
