19:02
2026-08-19
promptcube3.com
artificial-intelligence
Cerebras WSE-3 smokes H100 on Llama 3 70B inference at a
Cerebras Systems claims its Wafer-Scale Engine 3 (WSE-3) delivers 1,800 tokens per second at 23 kW for Llama 3 70B inference, versus about 850 tokens per second at 56 kW for an 8×H100 DGX system, yiel…