cd /news/large-language-models/deepseek-releases-v4-1-flash-says-it… · home topics large-language-models article
[ARTICLE · art-126350] src=siliconangle.com ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

DeepSeek releases V4.1-Flash, says it outperforms flagship V4-Pro

Hangzhou DeepSeek Artificial Intelligence Basic Technology Research Co. Ltd. released DeepSeek-V4.1-Flash, a 552-billion-parameter mixture-of-experts model it says beats its larger DeepSeek-V4-Pro on performance, cost, speed and total runtime. Starting Sept. 14, V4-Pro API requests will be answered by V4.1-Flash and billed at the smaller model's rates, cutting output costs roughly 70% from $3.96 to $1.20 per million tokens at peak, until a V4.1-Pro version launches. DeepSeek's benchmark table puts V4.1-Flash at 90.6 on Terminal-Bench 2.1, ahead of Anthropic's Claude Opus 5 at 89.1 and OpenAI's GPT-5.6 Sol at 88.8, though both U.S. models still lead on GPQA Diamond; weights are on Hugging Face under the MIT license.

by read3 min views7 publishedSep 10, 2026
DeepSeek releases V4.1-Flash, says it outperforms flagship V4-Pro
Image: Siliconangle (auto-discovered)

DeepSeek releases V4.1-Flash, says it outperforms flagship V4-Pro

Chinese artificial intelligence startup Hangzhou DeepSeek Artificial Intelligence Basic Technology Research Co. Ltd. today released DeepSeek-V4.1-Flash, the smallest model in a new architecture family. The company said tests by multiple parties put the open-weight model ahead of its much larger DeepSeek-V4-Pro on performance, cost, speed and total runtime.

Starting Sept. 14, requests sent to V4-Pro through DeepSeek’s application programming interface will be answered by V4.1-Flash and billed at the smaller model’s rates until a V4.1-Pro version launches. V4-Flash and the experimental vision model DeepSeek shipped in August are both retired. Calls to either now land on V4.1-Flash.

V4.1-Flash is a mixture-of-experts model with 552 billion parameters, close to double the 284 billion in V4-Flash. A new causal encoder-decoder design keeps just 8 billion parameters active while the model processes a prompt and 16 billion while it generates output. Image understanding, offered only in that experimental release last month, is now built into the model itself.

Much of the engineering went into shrinking the key-value cache. According to DeepSeek’s technical report, the model stores those entries in a four-bit floating-point format, and its global footprint comes to 890 bytes per token, about a quarter of what V4-Flash needs. Persistent cache storage on SSDs drops to roughly an eighth of the previous generation’s.

DeepSeek’s own benchmark table compares the model at maximum reasoning effort against Anthropic PBC’s Claude Opus 5 and OpenAI Group PBC’s GPT-5.6 Sol. V4.1-Flash scored 90.6 on Terminal-Bench 2.1, narrowly ahead of Opus 5 at 89.1 and GPT-5.6 Sol at 88.8. On the DeepSWE v1.1 software engineering test it resolved 74.2% of tasks, compared with 74% for Opus 5 and 62.7% for V4-Pro. Both U.S. models still lead on the GPQA Diamond science reasoning benchmark.

Off-peak API pricing is 15 cents per million uncached input tokens and 60 cents per million output tokens, with rates doubling during weekday peak windows. Developers still calling V4-Pro pay $3.96 per million output tokens at peak, compared with $1.20 for V4.1-Flash, so the rerouting works out to a cut of roughly 70% on output.

Weights are available on Hugging Face under the MIT license, and the model is live in DeepSeek’s web and mobile apps. DeepSeek said it will work with the open-source community on inference support and explore further deployment options.

The launch comes the same day Anthropic named DeepSeek in its latest threat intelligence report as one of seven China-based labs it says ran distillation campaigns against Claude. Anthropic attributed more than 12.1 million exchanges over 14 days in July to DeepSeek.

DeepSeek grew out of Chinese hedge fund High-Flyer. Founder Liang Wenfeng reportedly contributed $3 billion to a funding round of more than $7.4 billion in June that valued the company above $50 billion.

Image: DeepSeek

Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.

  • 15M+ viewers of theCUBE videos , powering conversations across AI, cloud, cybersecurity and more
  • 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network

Are you an AWS customer?  Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/

About SiliconANGLE Media

SiliconANGLE,

theCUBE Network,

theCUBE Research,

CUBE365,

theCUBE AIand theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.

── more in #large-language-models 4 stories · sorted by recency
── more on @deepseek 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/deepseek-releases-v4…] indexed:0 read:3min 2026-09-10 ·