15:51
2026-06-23
pytorch.org
large-language-models
Serving DeepSeek-V4 on GB300 with SGLang: 5x Higher Throughput at the Same Interactivity Since Day-0
SGLang achieved a 5x throughput improvement for serving DeepSeek-V4 on NVIDIA GB300 hardware since its Day-0 launch, reaching ~11,200 tok/s/GPU at 50 tok/s/user through kernel and runtime optimization…