{"slug": "alibaba-releases-open-qwen3-8-with-2-4t-total-95b-active-parameters", "title": "Alibaba releases open Qwen3.8 with 2.4T total, 95B active parameters", "summary": "Alibaba released the open-source Qwen3.8-2.4T-A95B model weights on August 12, a 2.4-trillion-parameter sparse mixture-of-experts model with about 95 billion active parameters per token, making it the first Qwen-Max-class model released openly. The text-only checkpoint requires thinking mode, supports a 262,144-token native context extendable to 1,010,000 tokens, and is compatible with Hugging Face Transformers, vLLM, SGLang, and TokenSpeed. The hosted Qwen3.8-Max on Qwen Cloud adds vision input, non-thinking mode, and built-in tools, but independent benchmarks for the open checkpoint are not provided.", "body_md": "Developers and infrastructure teams can now download the model weights for Qwen3.8, the latest release from [Alibaba](https://home.alibabagroup.com/en-US/ir-corporate-governance-partnership?ref=runtimewire)'s large-model program. Alibaba published the [Qwen3.8-2.4T-A95B weights](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B?ref=runtimewire) on August 12. Qwen calls it the first Qwen-Max-class model released openly in its [Qwen3.8 announcement](https://qwen.ai/blog?id=qwen3.8&ref=runtimewire).\n\nThe repository provides Qwen3.8-2.4T-A95B as model weights and configuration files in Hugging Face Transformers format, with compatibility for vLLM, SGLang and [TokenSpeed](https://lightseek.org/tokenspeed/recipes/models?ref=runtimewire#qwen3-8). Qwen separately offers Qwen3.8-Max through [Qwen Cloud](https://www.qwencloud.com/models/qwen3.8-max?ref=runtimewire). According to Qwen's repository and cloud documentation, that hosted version supports vision input, non-thinking mode and built-in tools.\n\n### What Qwen released\n\nQwen's repository describes Qwen3.8-2.4T-A95B as a 2.4-trillion-parameter sparse mixture-of-experts model with about 95 billion parameters activated per token. Its 92 layers contain 512 experts, with 10 routed experts and one shared expert active at a time. That design increases the model's stored capacity without using all 2.4 trillion parameters for every inference step.\n\nThe repository says the open checkpoint is text-only, requires thinking mode and supports a native 262,144-token context that can be extended to 1,010,000 tokens. Its `reasoning_effort`\n\nsetting supports `xhigh`\n\nby default, `medium`\n\nand `low`\n\n. A separate `preserve_thinking`\n\ncontrol retains reasoning context from earlier messages.\n\nThe model card lists support for Hugging Face Transformers, [SGLang](https://docs.sglang.io/cookbook/autoregressive/Qwen/Qwen3.8?ref=runtimewire), [vLLM](https://recipes.vllm.ai/Qwen/Qwen3.8-2.4T-A95B?ref=runtimewire) and TokenSpeed.\n\nThis is far removed from the local-development case for smaller Qwen releases. RuntimeWire reported in June that [Qwen 3.6 27B had become practical for local coding work](/article/qwen-36-27b-local-development-piotr-migdal-quesma). Qwen3.8's large parameter count creates a substantial storage requirement before an operator accounts for serving memory, compute and inference throughput.\n\n### The open checkpoint is narrower than Qwen3.8-Max\n\nThe downloadable checkpoint is text-only and always operates in thinking mode. It does not accept image or video inputs, and developers cannot disable its reasoning mode through the documented interface.\n\nAccording to [Qwen's model repository](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B?ref=runtimewire), Qwen3.8-Max adds vision input, non-thinking mode, built-in tools and a default 1-million-token context window. The same repository directs users seeking managed, scalable inference without infrastructure maintenance to Qwen Cloud.\n\nThe two releases give developers a downloadable text model and a hosted Max product with additional input modes and tools.\n\n### The repository links to a model-specific license\n\nThe repository links to a [Qwen3.8-specific license file](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B/blob/main/LICENSE?ref=runtimewire). The supplied materials do not include the license text, so its commercial-use conditions cannot be summarized here.\n\n### What the published benchmarks establish\n\nThe supplied materials do not establish independent benchmark results for the downloadable Qwen3.8-2.4T-A95B checkpoint, while [Qwen's model card reports company-run comparisons for the hosted Qwen3.8-Max version](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B?ref=runtimewire). Those comparisons cover coding, agent and general-capability tests under varying harnesses, run counts and evaluation conditions. Performance and operating costs for third-party deployments of the downloadable checkpoint therefore remain unestablished by the supplied evidence.", "url": "https://wpnews.pro/news/alibaba-releases-open-qwen3-8-with-2-4t-total-95b-active-parameters", "canonical_source": "https://runtimewire.com/article/alibaba-qwen38-24t-open-weights-enterprise-license", "published_at": "2026-08-12 16:15:41+00:00", "updated_at": "2026-08-12 16:37:59.171620+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "ai-research"], "entities": ["Alibaba", "Qwen", "Qwen3.8-2.4T-A95B", "Qwen3.8-Max", "Qwen Cloud", "Hugging Face", "vLLM", "SGLang"], "alternates": {"html": "https://wpnews.pro/news/alibaba-releases-open-qwen3-8-with-2-4t-total-95b-active-parameters", "markdown": "https://wpnews.pro/news/alibaba-releases-open-qwen3-8-with-2-4t-total-95b-active-parameters.md", "text": "https://wpnews.pro/news/alibaba-releases-open-qwen3-8-with-2-4t-total-95b-active-parameters.txt", "jsonld": "https://wpnews.pro/news/alibaba-releases-open-qwen3-8-with-2-4t-total-95b-active-parameters.jsonld"}}