00:00
2026-09-22
g-ftech.com
ai-infrastructure
Pushing the Limits: Extreme Inference Speedup of Qwen 3.8 27B on NVIDIA B300 (100 to 10k+ tok/s)
A custom inference stack running Qwen 3.8 27B on NVIDIA B300 SXM6 GPUs reached 9,825 tokens per second on 8× B300 with DP4×TP2 plus Suffix4, and a maximum cluster ceiling of 16,931 tokens per second o…