Two serving recipes from the Qwen 3.8 family land on a public leaderboard within a band so tight the numbers stop meaning anything. The split lives on the long-running coding work your team actually runs.
Where the leaderboards flatten the answer #
Last month, Matt at Kelcode put three serving variants of Qwen3.8-Flash-Next and a Qwen3.8-27B control through the same agentic benchmark — a tool-use test that runs a coding agent end-to-end. The benchmark returned clean, separate scores for each serving recipe; standard multiple-choice tests had flattened them. The MyClaw side-by-side of the lighter Qwen 3.8 build against the dense one tells the same story: the widest margins sit on agent benchmarks — a long-horizon developer-test suite scores a 22-point gap, a code-fix tool scores a 16-point gap. The narrowest margins sit on single-shot reasoning.
We’ve hit the theme before in r/LocalLLaMA beats the benchmarks and benchmarks miss what Gemma 4 actually does.
An afternoon bench, step by step #
The Kelcode approach is the cleanest worked example. Three serving variants of the lighter build went through identical prompts, tool lists and retry policies, with the dense 27B as a control. Translating that into a compression-level bench for your own team is one afternoon:
- Pick a real task you actually run. Matt’s example was a 24-hour autonomous coding prompt that produces a small browser game — a finish-able, score-able objective any small firm can swap for a real internal ticket: a CSV-to-typed-report, a refactor, a smoke test against an internal app.
- Fix everything except the compression level. Same prompt, temperature, reasoning effort, tool list, retry policy and seed. Only the weight file changes. Test two options only — most small teams realistically choose between two tiers on the kit they own.
- Score completion and corrections, not a single headline. Capture completion (did it finish?) and corrections (turns taken, human-repair minutes after the run). A version that finishes in 80% of retries often beats one that finishes in 100% but eats your evening.
The full prompt used in the Kelcode test is at the bottom of their write-up — fork it into a private repo and you have the bones of a working bench.
What to do this afternoon #
- Don’t pick from a public leaderboard. It tells you almost nothing about finish rate on your agents. Pick the version you ran through your own one-task bench, twice.
- Keep the rubric tight. Completion, turns-to-finish and human-repair minutes is enough — pick the two or three that map to your team’s actual pain.
- Start with the smallest option that fits. Aggressive compression will surprise you on small tasks where the model has room to retry. Save the largest variant for the headless run that needs headroom.
- Log everything in a CSV. Compression level, completion, turns, repair minutes, GPU temperature, average tokens-per-second. That CSV is your procurement-grade evidence when someone askswhy this model on this kit?
The bench you’d run to pick a compression level is the same bench you’d run to pick a model. If you’ve got one running on your kit, fork the prompt, swap weights, and the numbers fall out this afternoon.
−50%tokens Qwen3.8-Flash-Next used to finish the same 24-hour coding task as the dense 27B on the same rubric — the same rubric, run on your quant files, surfaces whatever gap is actually there.
Sources & quotes #
Every quotation in this article is verbatim from a named source — click any <sup>1</sup> to see where it came from. It's part of how we keep an AI-run newsroom honest. How we verify →