# Qwen3.8-27B

> Source: <https://tokenstead.ai/models/qwen3-8-27b>
> Published: 2026-08-19 22:19:06+00:00

# Qwen3.8-27B

enthusiast**27B dense with a vision encoder**, built on the Qwen 3.5 architectural foundation. Hybrid attention layout - Gated DeltaNet (linear attention) layers interleaved with Gated Attention in a 3:1 pattern across 64 layers - plus a ~248K vocabulary and Multi-Token Prediction (MTP) training. Multimodal: text, images, and hour-scale video in; text out.

-
**Context:** 262,144 tokens native, extensible to ~1M via YaRN. -
**Reasoning and tools:** hybrid thinking - on by default, disable per request;`reasoning_effort`

(xhigh default / medium / low) and`preserve_thinking`

. Improved tool calling, including parsing of nested JSON objects, plus Developer Role support for agentic tools. -
**Family:** the open-weight sibling of cloud-only Qwen3.8-Max; also ships alongside the open Qwen3.8-2.4T-A95B (thinking-only).

**Open weights** under Apache-2.0 at `Qwen/Qwen3.8-27B`

- released 2026-08-05, with a v2 weights update on 2026-08-14. Unsloth had day-zero quant access: its Qwen3.8 quants passed 5.1M downloads in the first five days.

**Unsloth Dynamic 3.0 (2026-08-19).** The re-quant calibrates on a higher-quality imatrix dataset (agentic coding, chat, multilingual) and improves layer selection - Unsloth measures >10% better top-1 accuracy at the same file size than every other quant provider (Divergence-300 @32, KL Divergence). The MTP module is dropped from UD-Q2_K_XL and below, saving ~500MB. What fits where (sizes are exact HF file sizes):

-
**UD-Q4_K_XL**(17.56GB): the quality pick - ~19GB RAM, so 24GB-class VRAM (RTX 4090/5080) or any 24GB+ Mac. -
**UD-Q2_K_XL**(9.83GB): the Dynamic 3.0 headline quant - +8% top-1 accuracy vs the next-best provider at the same size. ~11GB RAM. -
**UD-IQ1_S**(6.19GB): 1-bit, 89% smaller than BF16, runs in 8GB RAM at ~72-77% top-1 agreement (72% per the docs, 77% in the launch announcement). -
**NVFP4:** vLLM on NVIDIA Blackwell only - ~1.5x faster than BF16 with 92-97% accuracy recovery, fits 24GB VRAM.

Run the GGUFs via `llama.cpp -hf`

or Unsloth Desktop; there is no verified local Ollama tag yet. **Honest framing:** “by far the strongest model for its size” is Unsloth’s characterization - Model2-Max 3B still beats it on ARC-Challenge and MMLU-Pro - and the Dynamic 3.0 deltas are vendor-reported until independently replicated.

- 27.0B
- 262k
- apache 2.0
- Aug 2026

## Scores

## Run it locally

Per-quant memory needs and a static "can you run it?" reference - no rig entry required

### Can you run it? - reference rigs

| Rig | UD-IQ1_S | UD-Q2_K_XL | NVFP4 | UD-Q4_K_XL |
|---|---|---|---|---|
| 4x H100 80GB (320GB) | fast 1190.2t/s | fast 749.6t/s |
|

[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.

## Download options

## Or run it in the cloud

No per-token API provider pricing tracked for Qwen3.8-27B yet.
For flagship list prices, see the
[calculator](/calculator).

[See who runs Alibaba in production →](/adoption/alibaba)

## Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.
