{"slug": "why-not-use-deepseek-flash-for-everything", "title": "Why not use Deepseek Flash for everything?", "summary": "DeepSeek V4.1-Flash now beats OpenAI's Sol on Brokk's internal coding benchmarks at a fraction of Sol's estimated model size, according to a Brokk blog post citing the DeepSWE v1.1 score-versus-average-cost-per-task chart dated September 22, 2026. The post attributes the shift to improved tooling harnesses, advances in reasoning and post-training, and refined distillation, and its author says DeepSeek Flash V4.1 has largely replaced most of his model usage despite costing more than Flash v4.0. The author also rates GLM 5.3 (Z.ai) as his go-to for about a month and calls Sol (OpenAI) good value but prone to overengineering at higher thinking levels.", "body_md": "# Why not use Deepseek Flash for everything?\n\nI started using AI heavily when a friend and Brokk founder [Jonathan Ellis](https://www.linkedin.com/in/jbellis/?ref=blog.brokk.ai) launched his own startup Brokk and I started using his then revolutionary tool [Brokk](https://github.com/BrokkAi/brokk-app?ref=blog.brokk.ai) in April of 2025, at that stage basically all models were pretty terrible and things like Cursor or Claude Code were basically useless (I like to call them tooling harnesses because they hook up tools like bash, read, write, internet search to models..and that is most of what they do), so this meant we had to use the latest and greatest to ever get anything done. Brokk despite being basically unusable for a mere mortal (lesson let a designer build the UX), it was light years ahead of Claude Code and Cursor and with the top end model we could get a ton of solid generated code out of things like GPT 4 and Sonnet 3.5. They routinely invented stuff, didnt know how to look into directories etc, what we knew about prompting then was limited and in general getting good work out of them required a lot of cleverness or at least a willingness to accept really terrible code and a drop in productivity (at least in single threaded modes of work).\n\nAt least with Brokk with the tools we had built and the effeciencies we engaged in to keep the context small (the total amount of information sent to the models) we could make good use of cheaper less high end models, but generally speaking the lesson was clear, if you could afford it always use the best, or else waste a lot more time and money getting bad results out of the cheaper stuff.\n\nFast forward to today and even cheap models can score well on difficult coding benchmarks (source: [https://deepswe.datacurve.ai/](https://deepswe.datacurve.ai/?ref=blog.brokk.ai)) and the results are way more about 'feel' than they are about any sort of qualitative result that one can measure easily.\n\n*DeepSWE v1.1 score versus average cost per task, September 22, 2026 ([source](https://deepswe.datacurve.ai/?ref=blog.brokk.ai)). Click the chart to view it at full size.*\n\nThis is because of a few things at once: the tooling harnesses of today all do the things Brokk was doing in 2025 and often with more quality and efficiency, the [advances](https://arxiv.org/abs/2501.12948?ref=blog.brokk.ai) [in](https://aclanthology.org/2026.findings-acl.1767/?ref=blog.brokk.ai) [reasoning](https://arxiv.org/abs/2607.22529?ref=blog.brokk.ai),the advances in [post training](https://arxiv.org/abs/2503.14476?ref=blog.brokk.ai), and [distillation becoming](https://arxiv.org/abs/2505.09388?ref=blog.brokk.ai) refined to the point that [Deepseek flash 4.1](https://www.deepseek.com/en/news/deepseek-v4-1-flash/?ref=blog.brokk.ai) which is a model that now beats Sol on our internal benchmarks and is a fraction of the estimated size of Sol.\n\n*DeepSeek V4.1-Flash benchmark comparison ([source](https://www.deepseek.com/en/news/deepseek-v4-1-flash/?ref=blog.brokk.ai)). Click the table to view it at full size.*\n\nFor my own experience I can tell you that:\n\n- Fable 5.1 (Anthropic) and Astra (OpenAI) are great general purpose models and I have used them a lot, but for many tasks they vastly overengineer a solution.\n- Sol (OpenAI) is overall a very good value but it also can overengineer on higher thinking levels and is sort of dumb on its default level (it works but you just need to prompt it a lot more). I do not think this is a bad choice honestly as your model for everything and between it and Luna really justifies the cheaper price.\n- GLM 5.3 (Z.ai) is very capable and has been my go to for about a month now.\n- GLM 5.3 Flash (Z.ai), Deepseek Flash v4 are both very cheap to use and for a large variety of straight forward coding or admin tasks are easily good enough.\n- Recently I have been using Deepseek Flash V4.1 and while more expensive to use than Flash v4.0 it has largely replaced most my model usage, and next month I do not plan on renewing my Anthropic, Z.ai or OpenAI subs at the maximum level as a result.\n\n## Conclusion\n\nI do not actually care or think it is super important which model you use anymore, use the model that fits your style and you can afford and after that really compared to any other point prior we get very good results from nearly any provider. Use Muse, use OpenAI, use Z.AI you should be basing these differences now based on who treats you will, respects your data or [does](https://blog.ferstar.org/en/posts/zcode-silent-workspace-snapshot-upload?ref=blog.brokk.ai) [not](https://thehackernews.com/2026/07/grok-build-uploads-entire-git.html?ref=blog.brokk.ai). For some gnarly hard problems I will still probably reach a lot for Astra and Fable, but I doubt I will be on a 20x Max sub anymore.", "url": "https://wpnews.pro/news/why-not-use-deepseek-flash-for-everything", "canonical_source": "https://blog.brokk.ai/why-not-use-deepseek-flash-for-everything/", "published_at": "2026-09-22 09:47:37+00:00", "updated_at": "2026-09-22 09:53:39.583702+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-tools", "developer-tools"], "entities": ["DeepSeek", "DeepSeek V4.1-Flash", "Brokk", "Jonathan Ellis", "OpenAI", "Sol", "Z.ai", "GLM 5.3"], "alternates": {"html": "https://wpnews.pro/news/why-not-use-deepseek-flash-for-everything", "markdown": "https://wpnews.pro/news/why-not-use-deepseek-flash-for-everything.md", "text": "https://wpnews.pro/news/why-not-use-deepseek-flash-for-everything.txt", "jsonld": "https://wpnews.pro/news/why-not-use-deepseek-flash-for-everything.jsonld"}}