{"slug": "qwen3-8-27b", "title": "Qwen3.8-27B", "summary": "Alibaba released the open-weight Qwen3.8-27B multimodal model on 2026-08-05 under Apache-2.0, featuring a 27B dense architecture with a vision encoder, 262,144-token native context, and hybrid thinking; Unsloth's Dynamic 3.0 quants for the model surpassed 5.1 million downloads in the first five days, with the UD-Q2_K_XL quant claiming +8% top-1 accuracy over competitors at the same size.", "body_md": "# Qwen3.8-27B\n\nenthusiast**27B dense with a vision encoder**, built on the Qwen 3.5 architectural foundation. Hybrid attention layout - Gated DeltaNet (linear attention) layers interleaved with Gated Attention in a 3:1 pattern across 64 layers - plus a ~248K vocabulary and Multi-Token Prediction (MTP) training. Multimodal: text, images, and hour-scale video in; text out.\n\n-\n**Context:** 262,144 tokens native, extensible to ~1M via YaRN. -\n**Reasoning and tools:** hybrid thinking - on by default, disable per request;`reasoning_effort`\n\n(xhigh default / medium / low) and`preserve_thinking`\n\n. Improved tool calling, including parsing of nested JSON objects, plus Developer Role support for agentic tools. -\n**Family:** the open-weight sibling of cloud-only Qwen3.8-Max; also ships alongside the open Qwen3.8-2.4T-A95B (thinking-only).\n\n**Open weights** under Apache-2.0 at `Qwen/Qwen3.8-27B`\n\n- released 2026-08-05, with a v2 weights update on 2026-08-14. Unsloth had day-zero quant access: its Qwen3.8 quants passed 5.1M downloads in the first five days.\n\n**Unsloth Dynamic 3.0 (2026-08-19).** The re-quant calibrates on a higher-quality imatrix dataset (agentic coding, chat, multilingual) and improves layer selection - Unsloth measures >10% better top-1 accuracy at the same file size than every other quant provider (Divergence-300 @32, KL Divergence). The MTP module is dropped from UD-Q2_K_XL and below, saving ~500MB. What fits where (sizes are exact HF file sizes):\n\n-\n**UD-Q4_K_XL**(17.56GB): the quality pick - ~19GB RAM, so 24GB-class VRAM (RTX 4090/5080) or any 24GB+ Mac. -\n**UD-Q2_K_XL**(9.83GB): the Dynamic 3.0 headline quant - +8% top-1 accuracy vs the next-best provider at the same size. ~11GB RAM. -\n**UD-IQ1_S**(6.19GB): 1-bit, 89% smaller than BF16, runs in 8GB RAM at ~72-77% top-1 agreement (72% per the docs, 77% in the launch announcement). -\n**NVFP4:** vLLM on NVIDIA Blackwell only - ~1.5x faster than BF16 with 92-97% accuracy recovery, fits 24GB VRAM.\n\nRun the GGUFs via `llama.cpp -hf`\n\nor Unsloth Desktop; there is no verified local Ollama tag yet. **Honest framing:** “by far the strongest model for its size” is Unsloth’s characterization - Model2-Max 3B still beats it on ARC-Challenge and MMLU-Pro - and the Dynamic 3.0 deltas are vendor-reported until independently replicated.\n\n- 27.0B\n- 262k\n- apache 2.0\n- Aug 2026\n\n## Scores\n\n## Run it locally\n\nPer-quant memory needs and a static \"can you run it?\" reference - no rig entry required\n\n### Can you run it? - reference rigs\n\n| Rig | UD-IQ1_S | UD-Q2_K_XL | NVFP4 | UD-Q4_K_XL |\n|---|---|---|---|---|\n| 4x H100 80GB (320GB) | fast 1190.2t/s | fast 749.6t/s |\n|\n\n[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.\n\n## Download options\n\n## Or run it in the cloud\n\nNo per-token API provider pricing tracked for Qwen3.8-27B yet.\nFor flagship list prices, see the\n[calculator](/calculator).\n\n[See who runs Alibaba in production →](/adoption/alibaba)\n\n## Inference cost over time\n\nData accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.", "url": "https://wpnews.pro/news/qwen3-8-27b", "canonical_source": "https://tokenstead.ai/models/qwen3-8-27b", "published_at": "2026-08-19 22:19:06+00:00", "updated_at": "2026-08-19 23:12:57.637387+00:00", "lang": "en", "topics": ["large-language-models", "generative-ai", "ai-products", "ai-tools", "artificial-intelligence"], "entities": ["Alibaba", "Qwen3.8-27B", "Unsloth", "Qwen3.8-Max", "Qwen3.8-2.4T-A95B", "Model2-Max 3B", "NVIDIA", "vLLM"], "alternates": {"html": "https://wpnews.pro/news/qwen3-8-27b", "markdown": "https://wpnews.pro/news/qwen3-8-27b.md", "text": "https://wpnews.pro/news/qwen3-8-27b.txt", "jsonld": "https://wpnews.pro/news/qwen3-8-27b.jsonld"}}