Developer’s V4.1 Flash launch highlights China’s AI battle, where efficiency and speed drive competition amid chip curbs
DeepSeekhas released its V4.1 Flash model, claiming it outperforms its previous flagship while cutting inference costs and boosting speeds – the latest salvo in China’s aggressive price-and-performance war.
The Chinese artificial intelligence developer said on Thursday that V4.1 Flash used a new “Causal-Encoder-Decoder” architecture. While built on a massive 552 billion-parameter framework, it relies on a Mixture-of-Experts (MoE) design.
In traditional AI, every query runs through the entire system. An MoE system routes tasks only to the specific subnetworks best suited for the job. By activating just 8 billion parameters to process inputs and 16 billion to generate responses, DeepSeek said it significantly reduced the computing power needed per request.
DeepSeek described V4.1 Flash as the smallest model in its new series, with native multimodal visual understanding. The model outperformed V4 Pro on benchmarks evaluating coding, cybersecurity and autonomous agent tasks, the company said.
pure-play AI developerslock horns in a fierce battle to commercialise AI across China.
With rising hardware costs and foreign chip export curbs tightening compute constraints, Chinese players are racing to offer efficient models that deliver high-end reasoning at fraction-of-a-cent operational costs.
On Terminal-Bench 2.1, which tests AI on real-world compute tasks, V4.1 Flash scored 90.6, surpassing OpenAI’s GPT-5.6 Sol at 88.8, Moonshot AI’s Kimi K3 at 88.3, and DeepSeek’s own V4 Pro at 87.9.