# Qwen3.8 27B Runs 200 tok/s on a Single RTX 5090

> Source: <https://twitter.com/sgl_project/status/2088281320422322413>
> Published: 2026-08-14 17:56:10+00:00

The king of small models is back! Qwen3.8-27B from

[@Alibaba_Qwen](https://x.com/Alibaba_Qwen)is open source, and Day-0 support is live in SGLang: - 206.1 tok/s decode on a single RTX 5090, with our NVFP4 plus DSpark - 38.28 tok/s decode on DGX Spark Qwen3.8-27B raises the bar again for what a small model can do on agentic planning and long-horizon tasks. Long live the (small model) king! Run it locally with SGLang 👇We promised open weights for Qwen3.8. Now, time to meet them! 🎉
⚡ Qwen3.8-27B:
- A native multimodal dense model. With just 27B parameters, it outperforms Qwen3.7-Plus overall and shines in real-world coding & office workflows.
- 262K native context, easily extendable to 1M
