cd /news/large-language-models/qwen3-8-27b · home topics large-language-models article
[ARTICLE · art-103642] src=tokenstead.ai ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Qwen3.8-27B

Alibaba released the open-weight Qwen3.8-27B multimodal model on 2026-08-05 under Apache-2.0, featuring a 27B dense architecture with a vision encoder, 262,144-token native context, and hybrid thinking; Unsloth's Dynamic 3.0 quants for the model surpassed 5.1 million downloads in the first five days, with the UD-Q2_K_XL quant claiming +8% top-1 accuracy over competitors at the same size.

read3 min views5 publishedAug 19, 2026
Qwen3.8-27B
Image: Tokenstead (auto-discovered)

enthusiast27B dense with a vision encoder, built on the Qwen 3.5 architectural foundation. Hybrid attention layout - Gated DeltaNet (linear attention) layers interleaved with Gated Attention in a 3:1 pattern across 64 layers - plus a ~248K vocabulary and Multi-Token Prediction (MTP) training. Multimodal: text, images, and hour-scale video in; text out.

Context: 262,144 tokens native, extensible to ~1M via YaRN. - Reasoning and tools: hybrid thinking - on by default, disable per request;reasoning_effort

(xhigh default / medium / low) andpreserve_thinking

. Improved tool calling, including parsing of nested JSON objects, plus Developer Role support for agentic tools. -

Family: the open-weight sibling of cloud-only Qwen3.8-Max; also ships alongside the open Qwen3.8-2.4T-A95B (thinking-only). Open weights under Apache-2.0 at Qwen/Qwen3.8-27B

  • released 2026-08-05, with a v2 weights update on 2026-08-14. Unsloth had day-zero quant access: its Qwen3.8 quants passed 5.1M downloads in the first five days.

Unsloth Dynamic 3.0 (2026-08-19). The re-quant calibrates on a higher-quality imatrix dataset (agentic coding, chat, multilingual) and improves layer selection - Unsloth measures >10% better top-1 accuracy at the same file size than every other quant provider (Divergence-300 @32, KL Divergence). The MTP module is dropped from UD-Q2_K_XL and below, saving ~500MB. What fits where (sizes are exact HF file sizes):

**UD-Q4_K_XL**(17.56GB): the quality pick - ~19GB RAM, so 24GB-class VRAM (RTX 4090/5080) or any 24GB+ Mac. -
**UD-Q2_K_XL**(9.83GB): the Dynamic 3.0 headline quant - +8% top-1 accuracy vs the next-best provider at the same size. ~11GB RAM. -
**UD-IQ1_S**(6.19GB): 1-bit, 89% smaller than BF16, runs in 8GB RAM at ~72-77% top-1 agreement (72% per the docs, 77% in the launch announcement). -

NVFP4: vLLM on NVIDIA Blackwell only - ~1.5x faster than BF16 with 92-97% accuracy recovery, fits 24GB VRAM.

Run the GGUFs via llama.cpp -hf

or Unsloth Desktop; there is no verified local Ollama tag yet. Honest framing: “by far the strongest model for its size” is Unsloth’s characterization - Model2-Max 3B still beats it on ARC-Challenge and MMLU-Pro - and the Dynamic 3.0 deltas are vendor-reported until independently replicated.

  • 27.0B
  • 262k
  • apache 2.0
  • Aug 2026

Scores #

Run it locally #

Per-quant memory needs and a static "can you run it?" reference - no rig entry required

Can you run it? - reference rigs

| Rig | UD-IQ1_S | UD-Q2_K_XL | NVFP4 | UD-Q4_K_XL |
|---|---|---|---|---|

| 4x H100 80GB (320GB) | fast 1190.2t/s | fast 749.6t/s | |

no -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudFit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.

Download options #

Or run it in the cloud #

No per-token API provider pricing tracked for Qwen3.8-27B yet.

For flagship list prices, see the
[calculator](/calculator).

[See who runs Alibaba in production →](/adoption/alibaba)

Inference cost over time #

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

── more in #large-language-models 4 stories · sorted by recency
── more on @alibaba 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/qwen3-8-27b] indexed:0 read:3min 2026-08-19 ·