The king of small models is back! Qwen3.8-27B from
@Alibaba_Qwenis open source, and Day-0 support is live in SGLang: - 206.1 tok/s decode on a single RTX 5090, with our NVFP4 plus DSpark - 38.28 tok/s decode on DGX Spark Qwen3.8-27B raises the bar again for what a small model can do on agentic planning and long-horizon tasks. Long live the (small model) king! Run it locally with SGLang 👇We promised open weights for Qwen3.8. Now, time to meet them! 🎉 ⚡ Qwen3.8-27B:
- A native multimodal dense model. With just 27B parameters, it outperforms Qwen3.7-Plus overall and shines in real-world coding & office workflows.
- 262K native context, easily extendable to 1M