cd /news/large-language-models/deepseek-launches-v4-1-flash-with-lo… · home topics large-language-models article
[ARTICLE · art-127636] src=techrepublic.com ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

DeepSeek Launches V4.1-Flash With Lower Memory and API Costs

DeepSeek launched V4.1-Flash on Thursday, a 552-billion-parameter mixture-of-experts model that activates roughly 8 billion parameters per input token and 16 billion per output token, with a one-million-token context window and API pricing as low as 0.02 yuan per million cached input tokens off-peak. DeepSeek says the model cuts key-value cache memory to one-quarter of the HBM and one-eighth of the SSD storage of the previous generation, with the cache footprint falling from 3,514 bytes per token to 890 bytes, and it scored 90.6 on Terminal-Bench 2.1 versus 88.8 for OpenAI's GPT-5.6 Sol, 88.3 for Moonshot AI's Kimi K3 and 87.9 for V4 Pro on DeepSeek's own benchmarks. DeepSeek released the model weights on Hugging Face under an MIT license, though the company notes the results come from its own testing and may not reflect independently reproduced conditions.

by read4 min views3 publishedSep 11, 2026
DeepSeek Launches V4.1-Flash With Lower Memory and API Costs
Image: Techrepublic (auto-discovered)

DeepSeek says its latest model lowers memory requirements and API costs while outperforming V4 Pro on several internal benchmarks.

Chinese AI company DeepSeek launched V4.1-Flash on Thursday as the smallest model in its new V4.1 architecture family, combining native visual understanding with an architecture designed to improve speed, throughput and serving costs.

The model has 552 billion total parameters in a mixture-of-experts (MoE) system, but DeepSeek says it activates approximately 8 billion parameters per input token and 16 billion per output token. DeepSeek says its new Causal Encoder-Decoder architecture, combined with new pretraining methods and larger-scale reinforcement learning, allows the model to deliver stronger results without using the full model for every request.

DeepSeek also gives V4.1-Flash a one-million-token context window, making its cache efficiency particularly important for long-running conversations and agent workloads.

In benchmark results published by DeepSeek, V4.1-Flash outperformed V4 Pro and several competing models on selected coding, cybersecurity and agent evaluations. On Terminal-Bench 2.1, it scored 90.6, compared with 88.8 for OpenAI’s GPT-5.6 Sol, 88.3 for Moonshot AI’s Kimi K3 and 87.9 for V4 Pro.

The model also scored 88.1 on Cybergym, ahead of V4 Pro at 83.3, Kimi K3 at 80 and GPT-5.6 Sol at 84.5. On DeepSWE v1.1, V4.1-Flash narrowly beat Anthropic’s Claude Opus 5, although the Anthropic model remained ahead on other evaluations.

Those results come from DeepSeek’s own testing and may not reflect performance under independently reproduced conditions or real-world workloads. DeepSeek has released the model weights on Hugging Face under an MIT license, allowing developers to conduct their own evaluations subject to the repository’s published terms.

The real change is in memory #

DeepSeek’s bigger selling point may be what happens outside the benchmark chart. The company says V4.1-Flash uses only one-quarter of the HBM and one-eighth of the SSD storage required for its previous generation’s key-value cache. SCMP reports that the cache footprint fell from 3,514 bytes per token in the previous Flash model to 890 bytes.

That matters for AI agents, where maintaining large amounts of context can become an expensive part of running millions of requests. Lower memory requirements could let businesses handle more concurrent workloads without simply throwing more hardware at the problem.

DeepSeek has also cut API pricing, saying off-peak cached input can cost as little as 0.02 yuan per million tokens. Actual costs will depend on usage time and the mix of cached input, uncached input and output tokens.

More must-read AI coverage

  •   [SS&C Intralinks DealCentre AI vs. Datasite: Which platform is built for the future of dealmaking?](https://www.techrepublic.com/article/dealcentre-ai-vs-datasite/)      
    
  •   [SS&C Intralinks FundCentre AI vs. Juniper Square: Which platform better supports modern private markets fund managers?](https://www.techrepublic.com/article/fundcentre-ai-vs-juniper-square/)      
    
  •   [Why Data, Not Models, Determines AI Success](https://www.techrepublic.com/sponsored/why-data-not-models-determines-ai-success/)      
    
    [The Rise of the AI-Native Factory: How Physical AI Is Transforming Manufacturing](https://www.eweek.com/a/artificial-intelligence/the-rise-of-the-ai-native-factory-how-physical-ai-is-transforming-manufacturing/)      

DeepSeek is retiring its own flagship #

The company is making an unusually aggressive move with its existing lineup. Starting Sept. 14, DeepSeek says V4 Pro requests will be routed to V4.1-Flash and charged at the Flash rate until V4.1-Pro launches.

“Given that the DeepSeek V4.1 Flash model comprehensively surpasses the V4 Pro in all metrics … it would not be appropriate to provide DeepSeek users with the originally underperforming V4 Pro model at a higher price,” Cui Tianyi, head of DeepSeek’s Harness team, said on X, according to SCMP.

Older V4-Flash and V4-Flash-Vision-Exp endpoints are also being retired and temporarily routed to V4.1-Flash.

Organizations using the affected endpoints should test V4.1-Flash before the routing change, particularly if their applications depend on consistent output formats, latency targets or an approved model version.

What this means for AI buyers #

DeepSeek is betting that AI customers care less about the size of a model than how much useful work they get for each dollar and how quickly they get it.

That puts pressure on competitors to improve not just model intelligence but the economics of running those models at scale. For companies building coding tools, search interfaces or autonomous agents, a model that can maintain long contexts while using substantially less memory could be more valuable than a larger model that wins isolated benchmarks.

V4.1-Flash’s lower pricing and cache requirements could make it attractive for high-volume coding, search and agent workloads, but DeepSeek’s reported gains may vary across real deployments. Organizations considering the model — or affected by the V4 Pro routing change — should compare its output quality, latency, compatibility and total infrastructure costs against their current deployments before switching.

── more in #large-language-models 4 stories · sorted by recency
── more on @deepseek 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/deepseek-launches-v4…] indexed:0 read:4min 2026-09-11 ·