cd /news/artificial-intelligence/alibabas-qwen3-8-model-goes-live-on-… · home topics artificial-intelligence article
[ARTICLE · art-94153] src=cryptobriefing.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Alibaba’s Qwen3.8 model goes live on Nvidia’s GB300, hits 4,000 tokens per second

Alibaba's Qwen team launched Qwen3.8-Max on August 3, a 2.4 trillion parameter sparse Mixture-of-Experts model that runs at over 4,000 tokens per second per GPU on Nvidia's GB300 NVL72 hardware, scoring 86.1 on OSWorld-Verified. The model, which handles text, images, and video with a 1 million token context window, is priced at $2 per million input tokens and $6 per million output tokens. Alibaba plans to release open weights for Qwen3.8-Max and a smaller Qwen3.8-27B variant within the week following launch.

read2 min views1 publishedAug 12, 2026
Alibaba’s Qwen3.8 model goes live on Nvidia’s GB300, hits 4,000 tokens per second
Image: Cryptobriefing (auto-discovered)

Via logodix.com

The 2.4 trillion parameter model represents Alibaba's most aggressive play yet in the global AI arms race, with open weights coming soon

Alibaba just put the global AI leaderboard on notice. The company’s Qwen team launched Qwen3.8-Max on August 3, a massive 2.4 trillion parameter model that runs at over 4,000 tokens per second per GPU on Nvidia’s GB300 NVL72 hardware.

To put that speed in perspective, 4,000 tokens per second per GPU means the model can generate roughly 3,000 words of text every single second on a single chip.

What Qwen3.8-Max actually is #

The model operates as a sparse Mixture-of-Experts system, a design pattern where only a fraction of the model’s total parameters activate for any given input. Of its 2.4 trillion total parameters, approximately 95 billion are active per token.

Qwen3.8-Max handles text, images, and video natively as a multimodal system. It also supports a context window of 1 million tokens, which translates to roughly 750,000 words of input.

On the OSWorld-Verified benchmark, a test designed to measure how well AI models can autonomously complete computer tasks, Qwen3.8-Max scored 86.1. That places it ahead of both Anthropic’s Claude Fable 5 (85.0) and OpenAI’s GPT-5.6 Sol Max. It ranks as the second highest-performing AI model overall behind Fable 5 on broader evaluations.

The model targets advanced coding, long-horizon planning tasks, and complex agentic workflows. Alibaba is pricing API access at $2 per million input tokens and $6 per million output tokens.

The hardware story matters just as much #

The GB300 NVL72 represents Nvidia’s latest rack-scale inference platform, and early indications suggest it can deliver over 4,500 tokens per second per GPU for Qwen family models under optimized conditions.

Open weights and the competitive landscape #

Alibaba plans to release open weights for both Qwen3.8-Max and a smaller variant, Qwen3.8-27B. Those were expected to become available in the week following launch, giving researchers, startups, and enterprise developers direct access to fine-tune and deploy the models.

At $2 per million input tokens, Qwen3.8-Max undercuts the premium pricing tiers of both Claude and GPT on a per-token basis while claiming comparable or superior performance on key benchmarks.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our

Editorial Policy.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @alibaba 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/alibabas-qwen3-8-mod…] indexed:0 read:2min 2026-08-12 ·