13:19
2026-10-02
runtimewire.com
ai-infrastructure
Cerebras says mixed-chip inference lifts throughput 5x in early results
Cerebras reported that disaggregating the prefill and decode stages of AI inference lifted throughput 5x in early results using the same number of its systems, without reducing token-generation speed,…