17:38
2026-08-25
lmsys.org
artificial-intelligence
Pushing the Limits of Serving DeepSeek-V4-Pro
DeepSeek-V4-Pro, a 1.6-trillion-parameter Mixture-of-Experts model, achieves 271 output tokens/s per node on H20 GPUs at batch size 1, compared with 383.7 tokens/s on B300, a 1.42× performance gap, ac…