Qwen3.8-Max scores 86.6 on Terminal-Bench 2.1, but the Opus 5 column is blank Alibaba released Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts model with 95 billion active parameters, claiming an 86.6 score on Terminal-Bench 2.1 and 86.1 on OSWorld-Verified, with the API live August 2 and the open-weight build Qwen3.8-2.4T-A95B landing August 13 under Apache 2.0. The comparison tables cited by Alibaba contain no Opus 5 figure on Terminal-Bench 2.1, and LLM Stats' independent head-to-head puts Opus 5.5 ahead of Qwen3.8 Max on composite score, 60.3 to 51.5, while Qwen is roughly 3.2x cheaper per token on a blended 3:1 input/output basis. The 2.4T-parameter open-weight release requires multi-node infrastructure rather than a single desk machine. 86.6. That’s the Terminal-Bench 2.1 score Alibaba is claiming for Qwen3.8-Max, and it’s why your feed lit up in August. Some of the coverage reads it as a sign the model “nears Opus 5.” Here is how close the evidence actually gets it. First, the part that matters most. Alibaba’s release announcement mirrored on OpenLM.ai https://openlm.ai/qwen3.8/ calls this the first time a Qwen-Max-class model ships with open weights. The model scales to 2.4 trillion parameters on the architecture foundation of Qwen 3.5. Per the shattered.io writeup https://shattered.io/qwen3-8-max-open-weights-benchmarks-2026/ , it’s a mixture-of-experts with 95 billion active parameters. The API went live August 2. The open-weight build, Qwen3.8-2.4T-A95B, landed August 13 under Apache 2.0. That license line is the real news. Everything else is a scoreboard. What the scoreboard actually says Read the comparison tables closely. One side-by-side lists 86.6 on Terminal-Bench 2.1 and 86.1 on OSWorld-Verified for Qwen. GPT-5.6 Sol Max gets 83.2 on OSWorld-Verified, so Qwen sits ahead of another frontier flagship on that agentic benchmark. That is the strongest case for “nears Opus 5”: on terminal and computer-use work, Qwen3.8-Max is posting numbers in frontier territory. The caveat is that the Opus 5 column in those tables is empty. Neither side-by-side we found lists an Opus 5 figure on Terminal-Bench 2.1, so 86.6 has no direct Opus 5 number beside it. The same writeup flags these as vendor numbers. It points to Claude Opus 5 $5/$25 per million tokens in/out and GPT-5.6 Sol $5/$30 https://andrew.ooo/answers/qwen3-8-max-vs-claude-opus-5-vs-gpt-5-6-sol-frontier-august-2026/ as the proven picks for production. I agree with that read. Qwen is being measured against the right tier. It just hasn’t been shown to match it. The closest independent head-to-head we found sets a limit on the claim. LLM Stats https://llm-stats.com/models/compare/claude-opus-5-5-vs-qwen3.8-max compares Opus 5.5 against Qwen3.8 Max. Opus 5.5 leads the composite score, 60.3 to 51.5. It wins all five benchmarks LLM Stats reports for both models. Qwen is roughly 3.2x cheaper per token on a blended 3:1 input/output basis. Different Opus version, different methodology. Still, the shape is clear. On agentic benchmarks, Qwen3.8-Max runs level with or ahead of GPT-5.6 Sol Max, which is a legitimate frontier result. On the independent composite, Opus still leads by a clear margin. Qwen is near the frontier, not yet at Opus, and it costs far less. How we’d treat a vendor claim We’ve been burned by single-source confidence. When we audit our own pipeline, we don’t trust one log. We cross-check the audit log, the git history and the worker logs before we call a fix confirmed. Three sources, one timeline. Do the same with a benchmark. A vendor leaderboard row is one source https://andrew.ooo/answers/qwen3-8-max-vs-claude-opus-5-vs-gpt-5-6-sol-frontier-august-2026/ . Before you swap it into an agent loop: - Run your own task set, not theirs. Terminal-Bench measures terminal work. Your repo has its own mess. - Check the harness. Agentic scores swing with the scaffold around the model. - Price the whole run, not the token. Cheaper tokens mean little if the model needs more retries. We learned a smaller version of this with Qwen3-VL https://www.gladlabs.io/posts/qwen3-vl-integration-gotchas-2e027bce . It’s a thinking model, and its reasoning trace shares the token budget with the answer. The spec sheet never mentions that. Production did. Can you actually run it? Not on a desk. Active parameters are 95 billion, but a mixture-of-experts still has to keep all 2.4 trillion weights somewhere reachable. That’s multi-node territory, or a hosted endpoint. If you wanted a local card-sized Qwen, look at the 27B release we covered https://www.gladlabs.io/posts/qwen-38-27b-needs-22000-tokens-to-draw-a-pelican-1a0e2558 instead. So who benefits from open weights at this size? Teams with a cluster and a compliance reason to keep data in-house. Inference providers who can host it. Researchers who want to poke at a frontier-class MoE without asking permission. Apache 2.0 means you can do all of that without a lawyer. For the rest of us, the weights matter indirectly. They put a price ceiling on closed models. Hosted Qwen at roughly a third of the per-token cost is a negotiating position, even if you never use it. Our take Don’t switch on the strength of 86.6. Don’t dismiss it either. Qwen3.8-Max is the strongest open-weight release we’ve seen from a vendor claiming frontier agentic numbers. On OSWorld-Verified it beats GPT-5.6 Sol Max, 86.1 to 83.2, and it prices far below the closed flagships. That puts it in the conversation with Opus. The independent comparison still has Opus 5.5 ahead, 60.3 to 51.5, and the comparison tables leave the Opus 5 column blank, so “nears” is the most the evidence supports. It’s also unproven where it counts: on your tasks, under your harness. Treat the benchmark as a reason to run a test, not a reason to rewrite your stack. Pull the weights if you have the iron. Hit the API if you don’t. Run your own evals for a week. Then decide with numbers you generated yourself. For the wider pattern, see our piece on open-source agents in autonomous workflows https://www.gladlabs.io/posts/the-expanding-role-of-open-source-llm-agents-in-au-c4e62c7c .