{"slug": "qwen3-8-27b-runs-200-tok-s-on-a-single-rtx-5090", "title": "Qwen3.8 27B Runs 200 tok/s on a Single RTX 5090", "summary": "Alibaba's Qwen3.8-27B open-source model achieves 206.1 tokens per second decode on a single Nvidia RTX 5090 using SGLang's NVFP4 and DSpark, with Day-0 support in SGLang. The 27B-parameter multimodal dense model outperforms Qwen3.7-Plus overall and supports 262K native context, extendable to 1M.", "body_md": "The king of small models is back! Qwen3.8-27B from\n\n[@Alibaba_Qwen](https://x.com/Alibaba_Qwen)is open source, and Day-0 support is live in SGLang: - 206.1 tok/s decode on a single RTX 5090, with our NVFP4 plus DSpark - 38.28 tok/s decode on DGX Spark Qwen3.8-27B raises the bar again for what a small model can do on agentic planning and long-horizon tasks. Long live the (small model) king! Run it locally with SGLang 👇We promised open weights for Qwen3.8. Now, time to meet them! 🎉\n⚡ Qwen3.8-27B:\n- A native multimodal dense model. With just 27B parameters, it outperforms Qwen3.7-Plus overall and shines in real-world coding & office workflows.\n- 262K native context, easily extendable to 1M", "url": "https://wpnews.pro/news/qwen3-8-27b-runs-200-tok-s-on-a-single-rtx-5090", "canonical_source": "https://twitter.com/sgl_project/status/2088281320422322413", "published_at": "2026-08-14 17:56:10+00:00", "updated_at": "2026-08-14 18:12:13.310177+00:00", "lang": "en", "topics": ["large-language-models", "generative-ai", "ai-products", "ai-infrastructure"], "entities": ["Alibaba", "Qwen3.8-27B", "SGLang", "Nvidia RTX 5090", "DGX Spark", "Qwen3.7-Plus"], "alternates": {"html": "https://wpnews.pro/news/qwen3-8-27b-runs-200-tok-s-on-a-single-rtx-5090", "markdown": "https://wpnews.pro/news/qwen3-8-27b-runs-200-tok-s-on-a-single-rtx-5090.md", "text": "https://wpnews.pro/news/qwen3-8-27b-runs-200-tok-s-on-a-single-rtx-5090.txt", "jsonld": "https://wpnews.pro/news/qwen3-8-27b-runs-200-tok-s-on-a-single-rtx-5090.jsonld"}}