{"slug": "cerebras-cs4", "title": "Cerebras CS4", "summary": "Cerebras Systems unveiled the CS-4, a rack-scale AI system powered by three WSE-3 Turbo wafers, delivering up to 30x faster inference than GPU systems and up to 10x more throughput per watt than its predecessor CS-3. The CS-4, built on the new Nexus Platform Architecture, reduces wafer-to-wafer interconnect latency to 2 microseconds, enabling over 1,000 tokens per second on models exceeding 10 trillion parameters, and cuts deployment time from days to hours.", "body_md": "**The Fastest AI**\n\nJust Got Faster.\n\nJust Got Faster.\n\nIntroducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers upto 30x faster inference compared to GPUs, enhanced economics, and a simple path todeploy hyperscale capacity. It is the architecture for frontier AI.\n\n### Three WSE-3 Turbo per System\n\nEach wafer delivers up to 2x the speed of the previous generation\n\n### More Performance per Wafer\n\nAll new power, cooling, and I/O unleashes even more performance per wafer\n\n### Nexus Rack-Sacle Platform\n\nEnables rapid deployment in hyperscale datacenters\n\n## Up to 30x faster than GPUs\n\nPowered by WSE-Turbo, CS-4 delivers up to 30x faster inference compared to GPU systems, setting a new record for the fastest inference available in production.\n\n## Higher ultrafast throughput\n\nThe CS-4 solution shifts the inference Pareto frontier, delivering up to 10x more throughput per watt than CS-3 while generating tokens up to 30x faster than production GPU systems. The result is a system designed to deliver both throughput and interactivity.\n\n## Frontier-ready architecture\n\nBy reducing wafer-to-wafer interconnect latency to 2 microseconds, CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters, preserving interactive decode performance at unprecedented scale.\n\n# BUILT FOR HYPERSCALE\n\nCS-4 is the first iteration of the newCerebras Nexus PlatformArchitecture. It is built around amodular concept with threefoundational elements: Compute,Power, and I/O – each withsignificant innovation to simplifymanufacturing, deployment,maintenance, and upgrades.\n\n## Modular compute backpack design\n\nCerebras has fundamentally re-imagined the server. Each Wafer-Scale Backpack is a self-contained assembly thatfolds the wafer, power conversion, direct liquid cooling, high-speed I/O, and control electronics into a compact 3D package with 50% fewer components. This design simplifies manufacturing and reduces deployment time from days to hours.\n\n## High-density power delivery\n\nWith power delivery just 0.5 millimeters away from the processor - roughly 100x closer than the roughly 50mm of conventional GPU boards - CS-4 nearly eliminates board-level power loss. This enables the delivery of twice as much power to the WSE-3T,enabling higher operating frequencies and faster token generation.\n\n## Next-gen wafer I/O interface\n\nCS-4 introduces a new programmable I/O subsystem that doubles I/O bandwidth and reduces latency,benefitting both aggregated and disaggregated solutions. The Wafer I/O Module also enables wafers to be linked within and across racks without a switch,for wafer-to-wafer latency as low as two microseconds that is key to interactivity for models with tens of trillions of parameters.\n\n## Deploy infrastructure then compute\n\nCS-4 separates the stable power, cooling, and network layer from its modular wafer-scale compute. The Cerebras PowerRack can be installed and facility-qualified before computearrives. Compute backpacks then slide into place and connect to power, cooling, and data—reducing deployment from days to hours while simplifying service and future upgrades at hyperscale.", "url": "https://wpnews.pro/news/cerebras-cs4", "canonical_source": "https://www.cerebras.ai/cs4", "published_at": "2026-08-19 00:28:18+00:00", "updated_at": "2026-08-19 00:40:47.967652+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-infrastructure", "ai-chips", "ai-products"], "entities": ["Cerebras Systems", "CS-4", "WSE-3 Turbo", "Nexus Platform", "CS-3"], "alternates": {"html": "https://wpnews.pro/news/cerebras-cs4", "markdown": "https://wpnews.pro/news/cerebras-cs4.md", "text": "https://wpnews.pro/news/cerebras-cs4.txt", "jsonld": "https://wpnews.pro/news/cerebras-cs4.jsonld"}}