{"slug": "deepseek-launches-v4-1-flash-with-lower-memory-and-api-costs", "title": "DeepSeek Launches V4.1-Flash With Lower Memory and API Costs", "summary": "DeepSeek launched V4.1-Flash on Thursday, a 552-billion-parameter mixture-of-experts model that activates roughly 8 billion parameters per input token and 16 billion per output token, with a one-million-token context window and API pricing as low as 0.02 yuan per million cached input tokens off-peak. DeepSeek says the model cuts key-value cache memory to one-quarter of the HBM and one-eighth of the SSD storage of the previous generation, with the cache footprint falling from 3,514 bytes per token to 890 bytes, and it scored 90.6 on Terminal-Bench 2.1 versus 88.8 for OpenAI's GPT-5.6 Sol, 88.3 for Moonshot AI's Kimi K3 and 87.9 for V4 Pro on DeepSeek's own benchmarks. DeepSeek released the model weights on Hugging Face under an MIT license, though the company notes the results come from its own testing and may not reflect independently reproduced conditions.", "body_md": "DeepSeek says its latest model lowers memory requirements and API costs while outperforming V4 Pro on several internal benchmarks.\n\nChinese AI company DeepSeek launched V4.1-Flash on Thursday as the smallest model in its new V4.1 architecture family, combining native visual understanding with an architecture designed to improve speed, throughput and serving costs.\n\nThe model has 552 billion total parameters in a mixture-of-experts (MoE) system, but DeepSeek says it activates approximately 8 billion parameters per input token and 16 billion per output token. DeepSeek says its new Causal Encoder-Decoder architecture, combined with new pretraining methods and larger-scale reinforcement learning, allows the model to deliver stronger results without using the full model for every request.\n\n[DeepSeek](https://www.techrepublic.com/article/deepseek-generative-ai-model-china/) also gives V4.1-Flash a one-million-token context window, making its cache efficiency particularly important for long-running conversations and agent workloads.\n\nIn benchmark results published by DeepSeek, V4.1-Flash outperformed V4 Pro and several competing models on selected coding, cybersecurity and agent evaluations. On Terminal-Bench 2.1, it scored 90.6, compared with 88.8 for OpenAI’s GPT-5.6 Sol, 88.3 for Moonshot AI’s Kimi K3 and 87.9 for V4 Pro.\n\nThe model also scored 88.1 on Cybergym, ahead of V4 Pro at 83.3, Kimi K3 at 80 and GPT-5.6 Sol at 84.5. On DeepSWE v1.1, V4.1-Flash narrowly beat [Anthropic’s Claude Opus 5](https://www.techrepublic.com/article/news-claude-opus-5-vending-bench-ai-agent-risks/), although the Anthropic model remained ahead on other evaluations.\n\nThose results come from [DeepSeek’s own testing](https://www.deepseek.com/en/news/deepseek-v4-1-flash/) and may not reflect performance under independently reproduced conditions or real-world workloads. DeepSeek has released the model weights on Hugging Face under an MIT license, allowing developers to conduct their own evaluations subject to the repository’s published terms.\n\n## The real change is in memory\n\nDeepSeek’s bigger selling point may be what happens outside the benchmark chart. The company says V4.1-Flash uses only one-quarter of the HBM and one-eighth of the SSD storage required for its previous generation’s key-value cache. [SCMP reports](https://www.scmp.com/tech/big-tech/article/3367051/deepseek-says-new-flash-ai-model-beats-kimi-k3-cyber-coding-benchmarks) that the cache footprint fell from 3,514 bytes per token in the previous Flash model to 890 bytes.\n\nThat matters for [AI agents](https://www.techrepublic.com/article/news-ai-agents-enterprise-security-gap/), where maintaining large amounts of context can become an expensive part of running millions of requests. Lower memory requirements could let businesses handle more concurrent workloads without simply throwing more hardware at the problem.\n\nDeepSeek has also cut API pricing, saying off-peak cached input can cost as little as 0.02 yuan per million tokens. Actual costs will depend on usage time and the mix of cached input, uncached input and output tokens.\n\n### More must-read AI coverage\n\n- \n        [SS&C Intralinks DealCentre AI vs. Datasite: Which platform is built for the future of dealmaking?](https://www.techrepublic.com/article/dealcentre-ai-vs-datasite/)      \n- \n        [SS&C Intralinks FundCentre AI vs. Juniper Square: Which platform better supports modern private markets fund managers?](https://www.techrepublic.com/article/fundcentre-ai-vs-juniper-square/)      \n- \n        [Why Data, Not Models, Determines AI Success](https://www.techrepublic.com/sponsored/why-data-not-models-determines-ai-success/)      \n- \n        [The Rise of the AI-Native Factory: How Physical AI Is Transforming Manufacturing](https://www.eweek.com/a/artificial-intelligence/the-rise-of-the-ai-native-factory-how-physical-ai-is-transforming-manufacturing/)      \n\n## DeepSeek is retiring its own flagship\n\nThe company is making an unusually aggressive move with its existing lineup. Starting Sept. 14, DeepSeek says V4 Pro requests will be routed to V4.1-Flash and charged at the Flash rate until V4.1-Pro launches.\n\n“Given that the DeepSeek V4.1 Flash model comprehensively surpasses the V4 Pro in all metrics … it would not be appropriate to provide DeepSeek users with the originally underperforming V4 Pro model at a higher price,” Cui Tianyi, head of DeepSeek’s Harness team, said on X, according to SCMP.\n\nOlder V4-Flash and V4-Flash-Vision-Exp endpoints are also being retired and temporarily routed to V4.1-Flash.\n\nOrganizations using the affected endpoints should test V4.1-Flash before the routing change, particularly if their applications depend on consistent output formats, latency targets or an approved model version.\n\n## What this means for AI buyers\n\nDeepSeek is betting that AI customers care less about the size of a model than how much useful work they get for each dollar and how quickly they get it.\n\nThat puts pressure on competitors to improve not just model intelligence but the economics of running those models at scale. For companies building coding tools, search interfaces or autonomous agents, a model that can maintain long contexts while using substantially less memory could be more valuable than a larger model that wins isolated benchmarks.\n\nV4.1-Flash’s lower pricing and cache requirements could make it attractive for high-volume coding, search and agent workloads, but DeepSeek’s reported gains may vary across real deployments. Organizations considering the model — or affected by the V4 Pro routing change — should compare its output quality, latency, compatibility and total infrastructure costs against their current deployments before switching.", "url": "https://wpnews.pro/news/deepseek-launches-v4-1-flash-with-lower-memory-and-api-costs", "canonical_source": "https://www.techrepublic.com/article/news-deepseek-v4-1-flash-costs-apac-china/", "published_at": "2026-09-11 18:10:32+00:00", "updated_at": "2026-09-12 13:11:05.118744+00:00", "lang": "en", "topics": ["large-language-models", "ai-products", "ai-infrastructure", "ai-agents", "ai-chips"], "entities": ["DeepSeek", "V4.1-Flash", "V4 Pro", "OpenAI", "GPT-5.6 Sol", "Moonshot AI", "Kimi K3", "Anthropic"], "alternates": {"html": "https://wpnews.pro/news/deepseek-launches-v4-1-flash-with-lower-memory-and-api-costs", "markdown": "https://wpnews.pro/news/deepseek-launches-v4-1-flash-with-lower-memory-and-api-costs.md", "text": "https://wpnews.pro/news/deepseek-launches-v4-1-flash-with-lower-memory-and-api-costs.txt", "jsonld": "https://wpnews.pro/news/deepseek-launches-v4-1-flash-with-lower-memory-and-api-costs.jsonld"}}