enthusiast27B dense with a vision encoder, built on the Qwen 3.5 architectural foundation. Hybrid attention layout - Gated DeltaNet (linear attention) layers interleaved with Gated Attention in a 3:1 pattern across 64 layers - plus a ~248K vocabulary and Multi-Token Prediction (MTP) training. Multimodal: text, images, and hour-scale video in; text out.
Context: 262,144 tokens native, extensible to ~1M via YaRN. -
Reasoning and tools: hybrid thinking - on by default, disable per request;reasoning_effort
(xhigh default / medium / low) andpreserve_thinking
. Improved tool calling, including parsing of nested JSON objects, plus Developer Role support for agentic tools. -
Family: the open-weight sibling of cloud-only Qwen3.8-Max; also ships alongside the open Qwen3.8-2.4T-A95B (thinking-only).
Open weights under Apache-2.0 at Qwen/Qwen3.8-27B
- released 2026-08-05, with a v2 weights update on 2026-08-14. Unsloth had day-zero quant access: its Qwen3.8 quants passed 5.1M downloads in the first five days.
Unsloth Dynamic 3.0 (2026-08-19). The re-quant calibrates on a higher-quality imatrix dataset (agentic coding, chat, multilingual) and improves layer selection - Unsloth measures >10% better top-1 accuracy at the same file size than every other quant provider (Divergence-300 @32, KL Divergence). The MTP module is dropped from UD-Q2_K_XL and below, saving ~500MB. What fits where (sizes are exact HF file sizes):
**UD-Q4_K_XL**(17.56GB): the quality pick - ~19GB RAM, so 24GB-class VRAM (RTX 4090/5080) or any 24GB+ Mac. -
**UD-Q2_K_XL**(9.83GB): the Dynamic 3.0 headline quant - +8% top-1 accuracy vs the next-best provider at the same size. ~11GB RAM. -
**UD-IQ1_S**(6.19GB): 1-bit, 89% smaller than BF16, runs in 8GB RAM at ~72-77% top-1 agreement (72% per the docs, 77% in the launch announcement). -
NVFP4: vLLM on NVIDIA Blackwell only - ~1.5x faster than BF16 with 92-97% accuracy recovery, fits 24GB VRAM.
Run the GGUFs via llama.cpp -hf
or Unsloth Desktop; there is no verified local Ollama tag yet. Honest framing: “by far the strongest model for its size” is Unsloth’s characterization - Model2-Max 3B still beats it on ARC-Challenge and MMLU-Pro - and the Dynamic 3.0 deltas are vendor-reported until independently replicated.
- 27.0B
- 262k
- apache 2.0
- Aug 2026
Scores #
Run it locally #
Per-quant memory needs and a static "can you run it?" reference - no rig entry required
Can you run it? - reference rigs
| Rig | UD-IQ1_S | UD-Q2_K_XL | NVFP4 | UD-Q4_K_XL |
|---|---|---|---|---|
| 4x H100 80GB (320GB) | fast 1190.2t/s | fast 749.6t/s | |
no -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudFit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.
Download options #
Or run it in the cloud #
No per-token API provider pricing tracked for Qwen3.8-27B yet.
For flagship list prices, see the
[calculator](/calculator).
[See who runs Alibaba in production →](/adoption/alibaba)
Inference cost over time #
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.