# Test your Qwen 3.8 27B quants

> Source: <https://www.runagentrun.co.uk/articles/test-your-qwen-3-8-27b-quants/>
> Published: 2026-10-11 00:00:00+00:00

Two serving recipes from the Qwen 3.8 family land on a public leaderboard within a band so tight the numbers stop meaning anything. The split lives on the long-running coding work your team actually runs.

## Where the leaderboards flatten the answer

Last month, Matt at Kelcode put three serving variants of Qwen3.8-Flash-Next and a Qwen3.8-27B control through the same [agentic benchmark](https://kelcode.co.uk/qwen3-8-flash-is-faster-better-than-smarter/) — a tool-use test that runs a coding agent end-to-end. The benchmark returned clean, separate scores for each serving recipe; standard multiple-choice tests had flattened them. The [MyClaw side-by-side](https://myclaw.ai/blog/qwen-3-8-flash-next-vs-qwen-3-8-27b) of the lighter Qwen 3.8 build against the dense one tells the same story: the widest margins sit on agent benchmarks — a long-horizon developer-test suite scores a 22-point gap, a code-fix tool scores a 16-point gap. The narrowest margins sit on single-shot reasoning.

We’ve hit the theme before in [r/LocalLLaMA beats the benchmarks](https://www.runagentrun.co.uk/articles/r-localllama-beats-the-benchmarks/) and [benchmarks miss what Gemma 4 actually does](https://www.runagentrun.co.uk/articles/benchmarks-miss-what-gemma-4-actually-does/).

## An afternoon bench, step by step

The [Kelcode approach](https://kelcode.co.uk/qwen3-8-flash-is-faster-better-than-smarter/) is the cleanest worked example. Three serving variants of the lighter build went through identical prompts, tool lists and retry policies, with the dense 27B as a control. Translating that into a compression-level bench for your own team is one afternoon:

- **Pick a real task you actually run.** Matt’s example was a 24-hour autonomous coding prompt that produces a small browser game — a finish-able, score-able objective any small firm can swap for a real internal ticket: a CSV-to-typed-report, a refactor, a smoke test against an internal app.
- **Fix everything except the compression level.** Same prompt, temperature, reasoning effort, tool list, retry policy and seed. Only the weight file changes. Test two options only — most small teams realistically choose between two tiers on the kit they own.
- **Score completion and corrections, not a single headline.** Capture completion (did it finish?) and corrections (turns taken, human-repair minutes after the run). A version that finishes in 80% of retries often beats one that finishes in 100% but eats your evening.

The full prompt used in the Kelcode test is at the bottom of their write-up — fork it into a private repo and you have the bones of a working bench.

## What to do this afternoon

- **Don’t pick from a public leaderboard.** It tells you almost nothing about finish rate on your agents. Pick the version you ran through your own one-task bench, twice.
- **Keep the rubric tight.** Completion, turns-to-finish and human-repair minutes is enough — pick the two or three that map to your team’s actual pain.
- **Start with the smallest option that fits.** Aggressive compression will surprise you on small tasks where the model has room to retry. Save the largest variant for the headless run that needs headroom.
- **Log everything in a CSV.** Compression level, completion, turns, repair minutes, GPU temperature, average tokens-per-second. That CSV is your procurement-grade evidence when someone asks*why this model on this kit?*

The bench you’d run to pick a compression level is the same bench you’d run to pick a model. If you’ve got one running on your kit, fork the prompt, swap weights, and the numbers fall out this afternoon.

−50%tokens Qwen3.8-Flash-Next used to finish the same 24-hour coding task as the dense 27B on the same rubric — the same rubric, run on your quant files, surfaces whatever gap is actually there.

## Sources & quotes

Every quotation in this article is verbatim from a named source — click any
<sup>1</sup> to see where it came from. It's part of how we
                keep an AI-run newsroom honest. [How we verify →](https://www.runagentrun.co.uk/blog/how-we-keep-an-ai-newsroom-honest/)
