Qwen3.8 27B Runs 200 tok/s on a Single RTX 5090 Alibaba's Qwen3.8-27B open-source model achieves 206.1 tokens per second decode on a single Nvidia RTX 5090 using SGLang's NVFP4 and DSpark, with Day-0 support in SGLang. The 27B-parameter multimodal dense model outperforms Qwen3.7-Plus overall and supports 262K native context, extendable to 1M. The king of small models is back Qwen3.8-27B from @Alibaba Qwen https://x.com/Alibaba Qwen is open source, and Day-0 support is live in SGLang: - 206.1 tok/s decode on a single RTX 5090, with our NVFP4 plus DSpark - 38.28 tok/s decode on DGX Spark Qwen3.8-27B raises the bar again for what a small model can do on agentic planning and long-horizon tasks. Long live the small model king Run it locally with SGLang 👇We promised open weights for Qwen3.8. Now, time to meet them 🎉 ⚡ Qwen3.8-27B: - A native multimodal dense model. With just 27B parameters, it outperforms Qwen3.7-Plus overall and shines in real-world coding & office workflows. - 262K native context, easily extendable to 1M