Alibaba Releases Qwen 3.8 Max, Beats GPT 5.6 Sol And Fable On Many Benchmarks Alibaba released Qwen3.8-Max, a 2.4-trillion-parameter model with 95 billion active parameters, claiming it outperforms OpenAI's GPT-5.6 Sol and Anthropic's Claude Fable 5 on several benchmarks including PaperBench (93.0 vs 90.5) and IFBench (82.8 vs 72.7 and 63.5). The company said open weights will be available next week, marking the first time a Qwen-Max-class model is released publicly. Alibaba also open-sourced the smaller Qwen3.8-27B and highlighted the model's autonomous capabilities, including a 10-day coding run and a 125-hour research reproduction that beat the original paper's results by 2.7 points on AIME24. More and more Chinese labs are knocking on the doors of the frontier of AI. Alibaba has released Qwen3.8-Max, the newest and largest model in its Qwen family, and the company is positioning it directly against the best that OpenAI and Anthropic currently have on the market. The model scales up to 2.4 trillion parameters, with 95 billion active at any given time, and Alibaba says open weights are coming next week, marking the first time a Qwen-Max-class model will be released publicly rather than kept behind an API. That open-weight promise matters because Alibaba’s biggest models have typically stayed closed at launch. Qwen3.7-Max https://officechai.com/ai/qwen-3-7-max-benchmarks/ and the Qwen3.8-Max preview https://officechai.com/ai/alibaba-qwen-3-8/ that came before this release both launched as proprietary, cloud-hosted options only. Alibaba is now reversing that pattern for its top-tier model, following the same playbook it has used for smaller releases like Qwen3.6-35B-A3B https://officechai.com/ai/qwen3-6-35b-a3b-benchmarks/ , which shipped under Apache 2.0 from day one. Alongside Qwen3.8-Max, Alibaba is also open-sourcing a smaller sibling, Qwen3.8-27B, aimed at developers who want something they can run without a data center. Much of Alibaba’s pitch centers on autonomy over long stretches of work rather than single-shot answers. The company points to a 10-day autonomous coding run in which the model built a GitHub project called oh-my-cli from an empty folder, dispatching its own issues, running its own tests, and merging its own pull requests without a human reviewing each step. In a separate test, the model spent roughly 125 hours reproducing a published research paper on data selection for LLM training, then kept going on its own and landed on a method that beat the paper’s original results by 2.7 points on the AIME24 math benchmark. Alibaba also entered the model into a live Tianchi competition against 526 human teams and says it finished ahead of 458 of them within a 24-hour window. Qwen 3.8 Max Benchmarks Alibaba’s release includes comparisons against Claude Opus 4.8, Claude Fable 5, OpenAI’s GPT-5.6 Sol running at max reasoning, and its own predecessor, Qwen3.7-Max, across three groups of tests: coding agents, general agent work, and general capabilities. On coding, the picture is mixed but shows Qwen 3.8 Max is competitive with both GPT 5.6 Sol and Fable 5. Qwen3.8-Max scores 86.6 on Terminal Bench 2.1, behind Sol’s 88.8, but pulls ahead on PaperBench 93.0 vs 90.5 , AndroidBench 75.1 vs 74.0 , and Alibaba’s own QwenSWEBench 80.7 vs 73.5 and QwenQoderBench 58.4 vs 53.8 . Against Claude Fable 5, the results split the other way on the some software-engineering tests — Fable 5 leads on SWE-bench Pro 80.0 vs 67.7 , DeepSWE 1.1 70.0 vs 56.6 , and FrontierSWE 88.8 vs 73.5 — while Qwen3.8-Max edges ahead on PaperBench and QwenSVGBench. The general agent category, which covers multi-step office and workspace tasks, is where Qwen3.8-Max looks strongest relative to GPT-5.6 Sol. It beats Sol on CoWorkBench 74.8 vs 71.5 , WorkSpaceBench 67.7 vs 65.6 , JobBench 53.4 vs 45.4 , and WideSearch 81.9 vs a result Alibaba lists as unavailable for Sol on this test . Fable 5 stays ahead on most of these, including JobBench 57.4 and SkillsBench 70.9 vs Qwen’s 70.2 , though the gaps are narrower than on raw coding benchmarks. General capabilities tell a similar story of trade-offs rather than a clean sweep. Qwen3.8-Max posts a notably strong 82.8 on IFBench, ahead of both GPT-5.6 Sol 72.7 and Fable 5 63.5 , and leads on HealthBench 60.2 , PLawBench 73.2 , and PRBench-Finance 58.3 . It trails on Humanity’s Last Exam 43.6 vs Sol’s 47.2 and Fable’s 53.3 and on long-context recall tests like MRCR v2, where GPT-5.6 Sol’s 93.8 stays ahead of Qwen’s 92.9. Alibaba also ran the model through a 365-day e-commerce simulation it calls E-Commerce Bench, where the model manages a virtual retail operation starting with ¥100,000 in capital and has to handle supplier negotiations, seasonal demand, and even fraud detection against planted scam merchants. Qwen3.8-Max finished with a balance of ¥416,252, ahead of GLM 5.2 by 38 percent and more than double what the previous-generation Qwen3.7-Max managed on the same test. On a separate chip-design benchmark, the model autonomously optimized a cryptographic hardware accelerator over roughly 500 turns, cutting the design from 8,298 logic gates down to 678, and carried that reduction through to a physical layout that shrank the chip’s die area by 81 percent. Qwen 3.8 Max Pricing Qwen3.8-Max is available now through QwenCloud, priced at $2.00 per million input tokens and $6.00 per million output tokens, with implicit caching available at $0.25 per million tokens. That places it well below GPT-5.6 Sol and Claude Fable 5 on a per-token basis, continuing the pattern where Chinese labs undercut US frontier pricing even as they close the capability gap. Open weights for the full Qwen3.8-Max model are expected next week, alongside the smaller Qwen3.8-27B, which will let developers run comparable capability without going through Alibaba’s cloud at all. Pressure on Anthropic, OpenAI Qwen3.8-Max doesn’t arrive in isolation. It lands roughly two weeks after Moonshot AI’s Kimi K3 https://officechai.com/ai/kimi-k3-beats-fable-5-gpt-5-6-sol-on-frontend-code-arena/ , a 2.8-trillion-parameter open-weight model that already tops LMArena’s Frontend Code leaderboard ahead of Claude Fable 5, and it follows a string of DeepSeek releases that have spent the past year normalizing the idea that a Chinese lab can sit within a few points of the US frontier while giving away its weights for free. Between Kimi, DeepSeek, and now Qwen, there is no longer a single benchmark category where OpenAI and Anthropic hold a comfortable, undisputed lead — the gaps that remain are measured in single digits, not generations. For two companies preparing to test public markets on the strength of being the frontier, that is an uncomfortable backdrop. Anthropic and OpenAI have both filed confidentially for IPOs https://officechai.com/ai/anthropic-raises-65-billion-at-965-billion-valuation-is-now-worth-more-than-openai/ this year, with listings expected as early as the fall, and both are pitching investors on valuations built around the premise that their models are worth a premium few others can match. Every Chinese release that narrows that gap while undercutting them on price, and doing it in the open rather than behind an API, chips away at exactly the story they need Wall Street to believe.