08:21
2026-08-21
runtimewire.com
artificial-intelligence
SGLang publishes one-GPU Qwen3.8-27B recipes, claims 206.1 tokens per second
SGLang, the open-source inference project associated with Ying Sheng and Banghua Zhu, published deployment recipes for Alibaba's Qwen3.8-27B model, enabling single-GPU serving with NVFP4 quantization β¦