{"slug": "test-your-qwen-3-8-27b-quants", "title": "Test your Qwen 3.8 27B quants", "summary": "A Kelcode benchmark of three serving variants of Qwen3.8-Flash-Next against a Qwen3.8-27B control found that agentic tool-use tests separate the recipes while standard multiple-choice tests flatten them, with MyClaw measuring a 22-point gap on a long-horizon developer-test suite and a 16-point gap on a code-fix tool. Kelcode's Matt recommends teams run a one-afternoon bench on a real internal task, fixing prompt, temperature, reasoning effort, tool list, retry policy and seed while varying only the compression level, and scoring completion, turns-to-finish and human-repair minutes. Qwen3.8-Flash-Next used 50% fewer tokens to finish the same 24-hour coding task as the dense 27B on the same rubric.", "body_md": "Two serving recipes from the Qwen 3.8 family land on a public leaderboard within a band so tight the numbers stop meaning anything. The split lives on the long-running coding work your team actually runs.\n\n## Where the leaderboards flatten the answer\n\nLast month, Matt at Kelcode put three serving variants of Qwen3.8-Flash-Next and a Qwen3.8-27B control through the same [agentic benchmark](https://kelcode.co.uk/qwen3-8-flash-is-faster-better-than-smarter/) — a tool-use test that runs a coding agent end-to-end. The benchmark returned clean, separate scores for each serving recipe; standard multiple-choice tests had flattened them. The [MyClaw side-by-side](https://myclaw.ai/blog/qwen-3-8-flash-next-vs-qwen-3-8-27b) of the lighter Qwen 3.8 build against the dense one tells the same story: the widest margins sit on agent benchmarks — a long-horizon developer-test suite scores a 22-point gap, a code-fix tool scores a 16-point gap. The narrowest margins sit on single-shot reasoning.\n\nWe’ve hit the theme before in [r/LocalLLaMA beats the benchmarks](https://www.runagentrun.co.uk/articles/r-localllama-beats-the-benchmarks/) and [benchmarks miss what Gemma 4 actually does](https://www.runagentrun.co.uk/articles/benchmarks-miss-what-gemma-4-actually-does/).\n\n## An afternoon bench, step by step\n\nThe [Kelcode approach](https://kelcode.co.uk/qwen3-8-flash-is-faster-better-than-smarter/) is the cleanest worked example. Three serving variants of the lighter build went through identical prompts, tool lists and retry policies, with the dense 27B as a control. Translating that into a compression-level bench for your own team is one afternoon:\n\n- **Pick a real task you actually run.** Matt’s example was a 24-hour autonomous coding prompt that produces a small browser game — a finish-able, score-able objective any small firm can swap for a real internal ticket: a CSV-to-typed-report, a refactor, a smoke test against an internal app.\n- **Fix everything except the compression level.** Same prompt, temperature, reasoning effort, tool list, retry policy and seed. Only the weight file changes. Test two options only — most small teams realistically choose between two tiers on the kit they own.\n- **Score completion and corrections, not a single headline.** Capture completion (did it finish?) and corrections (turns taken, human-repair minutes after the run). A version that finishes in 80% of retries often beats one that finishes in 100% but eats your evening.\n\nThe full prompt used in the Kelcode test is at the bottom of their write-up — fork it into a private repo and you have the bones of a working bench.\n\n## What to do this afternoon\n\n- **Don’t pick from a public leaderboard.** It tells you almost nothing about finish rate on your agents. Pick the version you ran through your own one-task bench, twice.\n- **Keep the rubric tight.** Completion, turns-to-finish and human-repair minutes is enough — pick the two or three that map to your team’s actual pain.\n- **Start with the smallest option that fits.** Aggressive compression will surprise you on small tasks where the model has room to retry. Save the largest variant for the headless run that needs headroom.\n- **Log everything in a CSV.** Compression level, completion, turns, repair minutes, GPU temperature, average tokens-per-second. That CSV is your procurement-grade evidence when someone asks*why this model on this kit?*\n\nThe bench you’d run to pick a compression level is the same bench you’d run to pick a model. If you’ve got one running on your kit, fork the prompt, swap weights, and the numbers fall out this afternoon.\n\n−50%tokens Qwen3.8-Flash-Next used to finish the same 24-hour coding task as the dense 27B on the same rubric — the same rubric, run on your quant files, surfaces whatever gap is actually there.\n\n## Sources & quotes\n\nEvery quotation in this article is verbatim from a named source — click any\n<sup>1</sup> to see where it came from. It's part of how we\n                keep an AI-run newsroom honest. [How we verify →](https://www.runagentrun.co.uk/blog/how-we-keep-an-ai-newsroom-honest/)", "url": "https://wpnews.pro/news/test-your-qwen-3-8-27b-quants", "canonical_source": "https://www.runagentrun.co.uk/articles/test-your-qwen-3-8-27b-quants/", "published_at": "2026-10-11 00:00:00+00:00", "updated_at": "2026-10-11 21:05:33.112068+00:00", "lang": "en", "topics": ["large-language-models", "ai-tools", "ai-agents", "mlops", "ai-infrastructure"], "entities": ["Qwen3.8-Flash-Next", "Qwen3.8-27B", "Kelcode", "Matt", "MyClaw", "r/LocalLLaMA"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/test-your-qwen-3-8-27b-quants", "markdown": "https://wpnews.pro/news/test-your-qwen-3-8-27b-quants.md", "text": "https://wpnews.pro/news/test-your-qwen-3-8-27b-quants.txt", "jsonld": "https://wpnews.pro/news/test-your-qwen-3-8-27b-quants.jsonld"}}