Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72
Alibaba released the open weights for Qwen3.8-2.4T-A95B (Qwen3.8-Max), a 2.4T-parameter mixture-of-experts model with 95B activated parameters per token, designed for long-context reasoning and agenti…