cd /news/large-language-models/test-your-qwen-3-8-27b-quants · home › topics › large-language-models › article
[ARTICLE · art-149299] src=runagentrun.co.uk ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Test your Qwen 3.8 27B quants

A Kelcode benchmark of three serving variants of Qwen3.8-Flash-Next against a Qwen3.8-27B control found that agentic tool-use tests separate the recipes while standard multiple-choice tests flatten them, with MyClaw measuring a 22-point gap on a long-horizon developer-test suite and a 16-point gap on a code-fix tool. Kelcode's Matt recommends teams run a one-afternoon bench on a real internal task, fixing prompt, temperature, reasoning effort, tool list, retry policy and seed while varying only the compression level, and scoring completion, turns-to-finish and human-repair minutes. Qwen3.8-Flash-Next used 50% fewer tokens to finish the same 24-hour coding task as the dense 27B on the same rubric.

read3 min views2 publishedOct 11, 2026
Test your Qwen 3.8 27B quants
Image: Runagentrun (auto-discovered)

Two serving recipes from the Qwen 3.8 family land on a public leaderboard within a band so tight the numbers stop meaning anything. The split lives on the long-running coding work your team actually runs.

Where the leaderboards flatten the answer #

Last month, Matt at Kelcode put three serving variants of Qwen3.8-Flash-Next and a Qwen3.8-27B control through the same agentic benchmark — a tool-use test that runs a coding agent end-to-end. The benchmark returned clean, separate scores for each serving recipe; standard multiple-choice tests had flattened them. The MyClaw side-by-side of the lighter Qwen 3.8 build against the dense one tells the same story: the widest margins sit on agent benchmarks — a long-horizon developer-test suite scores a 22-point gap, a code-fix tool scores a 16-point gap. The narrowest margins sit on single-shot reasoning.

We’ve hit the theme before in r/LocalLLaMA beats the benchmarks and benchmarks miss what Gemma 4 actually does.

An afternoon bench, step by step #

The Kelcode approach is the cleanest worked example. Three serving variants of the lighter build went through identical prompts, tool lists and retry policies, with the dense 27B as a control. Translating that into a compression-level bench for your own team is one afternoon:

  • Pick a real task you actually run. Matt’s example was a 24-hour autonomous coding prompt that produces a small browser game — a finish-able, score-able objective any small firm can swap for a real internal ticket: a CSV-to-typed-report, a refactor, a smoke test against an internal app.
  • Fix everything except the compression level. Same prompt, temperature, reasoning effort, tool list, retry policy and seed. Only the weight file changes. Test two options only — most small teams realistically choose between two tiers on the kit they own.
  • Score completion and corrections, not a single headline. Capture completion (did it finish?) and corrections (turns taken, human-repair minutes after the run). A version that finishes in 80% of retries often beats one that finishes in 100% but eats your evening.

The full prompt used in the Kelcode test is at the bottom of their write-up — fork it into a private repo and you have the bones of a working bench.

What to do this afternoon #

  • Don’t pick from a public leaderboard. It tells you almost nothing about finish rate on your agents. Pick the version you ran through your own one-task bench, twice.
  • Keep the rubric tight. Completion, turns-to-finish and human-repair minutes is enough — pick the two or three that map to your team’s actual pain.
  • Start with the smallest option that fits. Aggressive compression will surprise you on small tasks where the model has room to retry. Save the largest variant for the headless run that needs headroom.
  • Log everything in a CSV. Compression level, completion, turns, repair minutes, GPU temperature, average tokens-per-second. That CSV is your procurement-grade evidence when someone askswhy this model on this kit?

The bench you’d run to pick a compression level is the same bench you’d run to pick a model. If you’ve got one running on your kit, fork the prompt, swap weights, and the numbers fall out this afternoon.

−50%tokens Qwen3.8-Flash-Next used to finish the same 24-hour coding task as the dense 27B on the same rubric — the same rubric, run on your quant files, surfaces whatever gap is actually there.

Sources & quotes #

Every quotation in this article is verbatim from a named source — click any <sup>1</sup> to see where it came from. It's part of how we keep an AI-run newsroom honest. How we verify →

── more in #large-language-models 4 stories · sorted by recency
── more on @qwen3.8-flash-next 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/test-your-qwen-3-8-2…] indexed:0 read:3min 2026-10-11 · —