Qwen opens a 2.4T model to self-hosting, with 95B active parameters On August 12th, Alibaba Cloud's Qwen team released Qwen3.8-2.4T-A95B, an open-weight mixture-of-experts model with 2.4 trillion total parameters and 95 billion active during inference, making it the first Qwen-Max-class model with downloadable weights. The release offers a self-hosted alternative to OpenAI's GPT-5.5 Pro, with adjustable reasoning depth and support for Hugging Face Transformers, vLLM, SGLang, and TokenSpeed, though running it requires substantial infrastructure. Qwen opens a 2.4T model to self-hosting, with 95B active parameters The Alibaba Cloud model activates 95B parameters per token, supports adjustable reasoning depth and offers a self-hosted alternative to OpenAI's GPT-5.5 Pro. By RuntimeWire Staff /author/runtimewire-staff ยท Published Primary source: GitHub Gist https://gist.github.com/wsxiaoys/e0286dc6bb624ff5fdf49e7f4c528ba3 Why it matters Qwen3.8 gives infrastructure teams downloadable weights for an Alibaba Max-class model with 95B active parameters, creating a self-hosted option for organizations that prioritize deployment control, data residency and inference customization over a fully managed API. On August 12th, Alibaba Cloud's Qwen team https://github.com/QwenLM/Qwen3.8?ref=runtimewire released Qwen3.8-2.4T-A95B /models/qwen/qwen3.8-2.4t-a95b , an open-weight mixture-of-experts model with 2.4 trillion total parameters and 95 billion active during inference. Developers can download its weights and configuration files rather than relying exclusively on a hosted model provider. The release extends Alibaba's push to put a Qwen-Max-class model into the hands of infrastructure teams willing to operate it. Alibaba https://www.sec.gov/Archives/edgar/data/1577552/000095017025090161/baba-20250331.htm?ref=runtimewire , founded in 1999 by Jack Ma and 17 co-founders, develops Qwen through its Alibaba Cloud business rather than through a separately financed startup. A very large model with local deployment options The Qwen3.8-2.4T-A95B model card https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B?ref=runtimewire describes a 92-layer architecture with 512 experts, 10 routed experts and one shared expert active in each mixture-of-experts block. Its native context length is 262,144 tokens, extensible to 1,010,000 tokens. Alibaba calls it the first Qwen-Max-class model released with open weights. The text-only model requires thinking mode, placing its reasoning inside