cd /news/large-language-models/qwen3-8-27b-runs-200-tok-s-on-a-sing… · home topics large-language-models article
[ARTICLE · art-97156] src=twitter.com ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Qwen3.8 27B Runs 200 tok/s on a Single RTX 5090

Alibaba's Qwen3.8-27B open-source model achieves 206.1 tokens per second decode on a single Nvidia RTX 5090 using SGLang's NVFP4 and DSpark, with Day-0 support in SGLang. The 27B-parameter multimodal dense model outperforms Qwen3.7-Plus overall and supports 262K native context, extendable to 1M.

read1 min views1 publishedAug 14, 2026
Qwen3.8 27B Runs 200 tok/s on a Single RTX 5090
Image: source

The king of small models is back! Qwen3.8-27B from

@Alibaba_Qwenis open source, and Day-0 support is live in SGLang: - 206.1 tok/s decode on a single RTX 5090, with our NVFP4 plus DSpark - 38.28 tok/s decode on DGX Spark Qwen3.8-27B raises the bar again for what a small model can do on agentic planning and long-horizon tasks. Long live the (small model) king! Run it locally with SGLang 👇We promised open weights for Qwen3.8. Now, time to meet them! 🎉 ⚡ Qwen3.8-27B:

  • A native multimodal dense model. With just 27B parameters, it outperforms Qwen3.7-Plus overall and shines in real-world coding & office workflows.
  • 262K native context, easily extendable to 1M
── more in #large-language-models 4 stories · sorted by recency
── more on @alibaba 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/qwen3-8-27b-runs-200…] indexed:0 read:1min 2026-08-14 ·