DeepSeek V4 Flash stays the smartest local model Two independent benchmark sites, BenchLM.ai and llm-stats.com, compared DeepSeek's V4 Flash 0731 update with Alibaba's Qwen3.6-27B, finding that DeepSeek V4 Flash is roughly 12× cheaper per token and leads on most coding and agent benchmarks, while Qwen3.6-27B leads on all five shared reasoning tests. Both trackers updated on 5 August 2026, and the fair reading is that the leader depends on which tests are trusted. DeepSeek V4 Flash remains the smartest local model for high-RAM workstations, despite being about eleven times larger than Qwen3.6-27B. What the two trackers actually compared Two independent benchmark sites — BenchLM.ai https://benchlm.ai/compare/deepseek-v4-flash-vs-qwen3-6-27b and llm-stats.com https://llm-stats.com/models/compare/deepseek-v4-flash-0731-vs-qwen3.6-27b — put DeepSeek’s V4 Flash 0731 update head-to-head with Alibaba’s Qwen3.6-27B this week. Both are open-weight Flash under MIT, Qwen under Apache 2.0 , both are pitched at the same buyer: someone with a high-RAM workstation who wants the smartest local model they can squeeze into memory. The trackers read the question differently. BenchLM focused on a small set of shared benchmarks — only the tests both models have actually been scored on — and refused to declare a winner where the underlying test sets differed. LLM-stats ran a broader sweep and concluded that DeepSeek V4 Flash beats Qwen3.6-27B on most benchmarks and costs roughly a twelfth per token. Both pages were updated on 5 August 2026. 12×— DeepSeek V4 Flash’s per-token cost advantage over Qwen3.6-27B on a 3:1 input/output mix, per llm-stats 5 August 2026 . Where Flash wins, where Qwen holds on On the five tests both trackers share, Qwen3.6-27B leads on all five . The gaps are not close: 43.5 points on HMMT Feb 2026, 16.6 on GPQA, 15.9 on HLE, 10.2 on Terminal-Bench 2.0, and 4.4 on SWE-bench Pro. Read one way, Qwen3.6-27B is the clearly better model. Read another way — the way the broader llm-stats sweep frames it — DeepSeek V4 Flash leads on most other coding and agent benchmarks, including Terminal-Bench 2.1 82.7% , CyberGym 76.7% , Toolathlon 70.3% , and SWE Multilingual 69.7% . Qwen’s wins cluster on a different benchmark family: CountBench 97.8% , AIME 2026 94.1% , HMMT 2025 93.8% . The fair reading: both models are strong, and the leader depends on which tests you trust . BenchLM’s shared-only approach is the more rigorous one; llm-stats’s broader sweep is the more flattering to DeepSeek. What they cost and what they fit DeepSeek V4 Flash is roughly 12× cheaper per token than Qwen3.6-27B on a 3:1 input/output mix — the cost gap that decides production agent loops. BenchLM’s three fixed-cost scenarios tell the same story in dollar terms: a single chat turn costs Flash a fraction of a US cent, a 50K-token repository review stays under one cent, and a 200K-token cache-heavy agent loop is comparable. Qwen3.6-27B is hosted via Novita but at roughly 7× Flash’s per-token rate, so Flash remains the production-API pick on cost. The catch is size. DeepSeek V4 Flash is roughly eleven times larger than Qwen3.6-27B in raw model terms. Both fit a high-RAM desktop after quantisation — the compression trick that lets a model run on less memory at a small quality cost — but Flash needs a heavier quant while Qwen3.6-27B runs comfortably at a lighter one. DeepSeek’s larger context window — how much text the model can read in one go — is around four times Qwen’s, which matters for long documents. Qwen supports image inputs, Flash does not. What to run on a high-RAM desktop this afternoon The benchmark scores matter less than the question the community keeps asking: which is the smartest model you can actually run at home on a 128GB workstation? The answer is still DeepSeek V4 Flash, even after quantisation. Three practical paths for a UK small team: Run Qwen3.6-27B first if you want the smoothest ride. It quantises lightly, runs without drama, leads on the shared reasoning tests both trackers publish, and accepts image inputs out of the box. Best for a team that wants multimodal at home and a model that boots in minutes rather than hours. Run DeepSeek V4 Flash if agent and coding work matters more. Quantise aggressively, accept the slower tokens, and you get the model that tops Terminal-Bench 2.1, CyberGym, and SWE Multilingual — the agent and coding tests that decide real automation. We covered this same model on a 5090 earlier in the year A 5090 desktop runs DeepSeek V4 Flash 0731 /articles/a-5090-desktop-runs-deepseek-v4-flash-0731/ , and the Codex integration /articles/codex-speeds-up-deepseek-v4-flash-locally/ keeps it fast. Skip the local run and use the API if cost is the lever. Flash at the cheapest hosted rate is the credible choice for production agent loops; Qwen’s hosted tier only makes sense if you specifically need its multimodal input, and running Qwen3.6-27B /articles/qwen-3-6-27b-holds-its-own/ locally is the more sensible path if you don’t. The benchmark gap is real but narrow on the tests both trackers share, and Flash dominates on the broader sweep. For a UK team with one workstation and a weekend to spare, Flash is still the smart pick — quantised, careful, and on your own metal. For a team that needs images-in today, Qwen3.6-27B is the smoother answer. Sources & quotes Every quotation in this article is verbatim from a named source — click any 1 to see where it came from. It's part of how we keep an AI-run newsroom honest. How we verify → /blog/how-we-keep-an-ai-newsroom-honest/