Qwen3.8-27B Alibaba released the open-weight Qwen3.8-27B multimodal model on 2026-08-05 under Apache-2.0, featuring a 27B dense architecture with a vision encoder, 262,144-token native context, and hybrid thinking; Unsloth's Dynamic 3.0 quants for the model surpassed 5.1 million downloads in the first five days, with the UD-Q2_K_XL quant claiming +8% top-1 accuracy over competitors at the same size. Qwen3.8-27B enthusiast 27B dense with a vision encoder , built on the Qwen 3.5 architectural foundation. Hybrid attention layout - Gated DeltaNet linear attention layers interleaved with Gated Attention in a 3:1 pattern across 64 layers - plus a ~248K vocabulary and Multi-Token Prediction MTP training. Multimodal: text, images, and hour-scale video in; text out. - Context: 262,144 tokens native, extensible to ~1M via YaRN. - Reasoning and tools: hybrid thinking - on by default, disable per request; reasoning effort xhigh default / medium / low and preserve thinking . Improved tool calling, including parsing of nested JSON objects, plus Developer Role support for agentic tools. - Family: the open-weight sibling of cloud-only Qwen3.8-Max; also ships alongside the open Qwen3.8-2.4T-A95B thinking-only . Open weights under Apache-2.0 at Qwen/Qwen3.8-27B - released 2026-08-05, with a v2 weights update on 2026-08-14. Unsloth had day-zero quant access: its Qwen3.8 quants passed 5.1M downloads in the first five days. Unsloth Dynamic 3.0 2026-08-19 . The re-quant calibrates on a higher-quality imatrix dataset agentic coding, chat, multilingual and improves layer selection - Unsloth measures 10% better top-1 accuracy at the same file size than every other quant provider Divergence-300 @32, KL Divergence . The MTP module is dropped from UD-Q2 K XL and below, saving ~500MB. What fits where sizes are exact HF file sizes : - UD-Q4 K XL 17.56GB : the quality pick - ~19GB RAM, so 24GB-class VRAM RTX 4090/5080 or any 24GB+ Mac. - UD-Q2 K XL 9.83GB : the Dynamic 3.0 headline quant - +8% top-1 accuracy vs the next-best provider at the same size. ~11GB RAM. - UD-IQ1 S 6.19GB : 1-bit, 89% smaller than BF16, runs in 8GB RAM at ~72-77% top-1 agreement 72% per the docs, 77% in the launch announcement . - NVFP4: vLLM on NVIDIA Blackwell only - ~1.5x faster than BF16 with 92-97% accuracy recovery, fits 24GB VRAM. Run the GGUFs via llama.cpp -hf or Unsloth Desktop; there is no verified local Ollama tag yet. Honest framing: “by far the strongest model for its size” is Unsloth’s characterization - Model2-Max 3B still beats it on ARC-Challenge and MMLU-Pro - and the Dynamic 3.0 deltas are vendor-reported until independently replicated. - 27.0B - 262k - apache 2.0 - Aug 2026 Scores Run it locally Per-quant memory needs and a static "can you run it?" reference - no rig entry required Can you run it? - reference rigs | Rig | UD-IQ1 S | UD-Q2 K XL | NVFP4 | UD-Q4 K XL | |---|---|---|---|---| | 4x H100 80GB 320GB | fast 1190.2t/s | fast 749.6t/s | | no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast =20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark. Download options Or run it in the cloud No per-token API provider pricing tracked for Qwen3.8-27B yet. For flagship list prices, see the calculator /calculator . See who runs Alibaba in production → /adoption/alibaba Inference cost over time Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.