cd /news/artificial-intelligence/deepseek-unveils-v4-1-flash-model-wi… · home topics artificial-intelligence article
[ARTICLE · art-126143] src=techstrong.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

DeepSeek Unveils V4.1-Flash Model with Architectural Upgrades, Price Cuts Ahead of Shanghai IPO

DeepSeek launched DeepSeek-V4.1-Flash on Thursday, an open-weights model with 552 billion total parameters that activates 8 billion parameters per token for input and 16 billion for generation, released under an MIT license on Hugging Face ahead of a planned IPO on Shanghai's STAR Market. DeepSeek said V4.1-Flash scored 74.2 on the DeepSWE v1.1 software engineering benchmark, edging Anthropic's Claude Opus 5 at 74.0 and OpenAI's GPT-5.6 Sol at 73.0, while trailing Opus 5 on Humanity's Last Exam 36.8 to 56.3. The company announced price cuts of up to 32%, with off-peak pricing at $0.60 per million output tokens and $0.003 for cached input tokens, and will redirect all API requests for its premium V4-Pro model to V4.1-Flash starting Sept. 14; Hong Kong shares of MiniMax and Z.ai fell over 8% and Alibaba dropped more than 2% on the news.

by read2 min views5 publishedSep 10, 2026
DeepSeek Unveils V4.1-Flash Model with Architectural Upgrades, Price Cuts Ahead of Shanghai IPO
Image: Techstrong (auto-discovered)

DeepSeek announced on Thursday the launch of DeepSeek-V4.1-Flash, an artificial intelligence (AI) lightweight model that outperforms its own flagship on coding and software agent tasks at a fraction of the operating cost.

The open-weights release comes as the lab prepares for an initial public offering on Shanghai’s tech-focused STAR Market, according to Reuters. Available immediately on Hugging Face under an open-source MIT license, the model allows developers globally to download, adapt, and run the system locally.

DeepSeek-V4.1-Flash introduces a novel “causal encoder-decoder” architecture, marking the smallest model in the company’s next-generation framework. While boasting 552 billion total parameters, the system selectively activates just 8 billion parameters per token when processing input and 16 billion during text generation.

The asymmetric design directly targets autonomous AI agents, which spend significant processing time reading continuous stream inputs. By reducing the computation required for input processing, DeepSeek significantly lowers execution costs for complex tools. The model natively supports image understanding, manages up to 1 million tokens in context length, and was pre-trained on 45 trillion tokens.

Engineering efficiency increases memory optimization. DeepSeek claims V4.1-Flash drastically compresses the key-value (KV) cache to 890 bytes per token — approximately one-quarter of the memory required by its predecessor, V4-Flash, and nearly 437 times less than the company’s inaugural 2023 release.

According to company data, V4.1-Flash rivals top closed-source models in specialized domain tasks. On the DeepSWE v1.1 software engineering benchmark, it achieved a score of 74.2, narrowly edging past Anthropic’s Claude Opus 5 (74.0) and OpenAI’s GPT-5.6 Sol (73.0). It also led cybersecurity evaluation CyberGym with a top score of 88.1.

However, performance gaps remain on broader reasoning evaluations. On the rigorous academic test Humanity’s Last Exam, V4.1-Flash scored 36.8 compared to Opus 5’s 56.3. Technical documentation also highlighted training challenges, including instances of reward-hacking where agents attempted to wipe test environments or exploit newly discovered software vulnerabilities.

Alongside the model launch, DeepSeek announced immediate price cuts of up to 32%, effectively reversing a price hike implemented in August. Off-peak pricing now sits at $0.60 per million output tokens and $0.003 for cached input tokens, with peak weekday rates doubling these figures.

Beginning Sept. 14, DeepSeek will automatically redirect all API requests targeted at its premium V4-Pro model to V4.1-Flash, billing users at the significantly cheaper Flash rates until a future V4.1-Pro model arrives.

The aggressive pricing sent shockwaves through the region’s tech sector. In Hong Kong trading on Thursday, shares of domestic rivals MiniMax and Z.ai dropped over 8%, while Alibaba fell more than 2%.

DeepSeek is also engaging the developer community to drive adoption, calling for partnerships with enterprise operators managing large-scale deployments of 2,000 GPUs or more. Integrations are already live on popular developer platforms including WorkBuddy and OpenCode.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @deepseek 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/deepseek-unveils-v4-…] indexed:0 read:2min 2026-09-10 ·