DeepSeek opens public beta API for V4-Flash as a three-way price war hits AI developers DeepSeek's V4-Flash, available since April 24, 2026, is priced at 1 yuan per million uncached input tokens, 2 yuan per million output tokens, and 0.02 yuan for cache-hit input, undercutting OpenAI's GPT-5.6 Luna at $0.20 input and $1.20 output per million tokens and Anthropic's Haiku 4.5 at $1 input and $5 output. The model, with 284 billion total parameters and 13 billion active per token, scores 79.0% on SWE-bench Verified and 91.6% on LiveCodeBench, according to DeepSeek's Hugging Face model card. Together AI raised $800 million at an $8.3 billion valuation on July 1, with annual bookings above $1.15 billion, and open-source model usage tripled over the past year, per TechCrunch. DeepSeek's V4-Flash is not a new July 31 launch, but its price still lands hard after OpenAI's July 30 cuts. If you're paying for high-volume AI calls, the comparison is now too large to ignore. The useful story here isn't that DeepSeek suddenly dropped a new public beta this morning. It didn't. DeepSeek's own API changelog lists V4-Pro and V4-Flash as available from April 24, 2026, and the Hugging Face model card gives the same release date. The fresher angle is more practical: DeepSeek's old deepseek-chat and deepseek-reasoner aliases were set to retire on July 24, while OpenAI has just cut GPT-5.6 prices. Developers now have to make a cost decision, not just read another benchmark thread. DeepSeek V4-Flash is still a serious entry in that decision. The official DeepSeek pricing page lists a one-million-token context window, a maximum output length of 384K tokens, OpenAI-format access at api.deepseek.com, and support for tool calls, JSON output and beta prefix completion. The model card lists 284 billion total parameters with 13 billion active per token through its MoE design. That's not a slogan. That's the shape of the product you would actually wire into an agent pipeline. The pricing is where this gets pointed. DeepSeek lists V4-Flash at 1 yuan per million uncached input tokens, 2 yuan per million output tokens and 0.02 yuan for cache-hit input. The dollar figures usually quoted around the model, roughly $0.14 input, $0.28 output and $0.003 cached input, are conversions of those yuan prices, not a separate U.S. tariff page. Either way, the gap is hard to miss. Axios reported that OpenAI cut GPT-5.6 Luna to $0.20 per million input tokens and $1.20 per million output tokens on July 30. Anthropic's own Claude pricing page lists Haiku 4.5 at $1 input and $5 output. DeepSeek undercuts Luna on input and crushes it on output. Price alone doesn't settle anything. You still need quality, reliability, data controls and procurement approval. But when output tokens cost less than a quarter of OpenAI's cheaper Luna tier, you don't get to wave the comparison away as a hobbyist talking point. The benchmark story is strong, but it needs clean sourcing The original draft leaned on developers testing the model on X and on a single unnamed developer saying Flash had surpassed V4-Pro-Preview. That's not good enough. A convenient social post with no verifiable record is exactly the sort of attribution that should come out before publication. The stronger numbers are already public. DeepSeek's Hugging Face model card reports V4-Flash Max at 79.0% on SWE-bench Verified, with V4-Pro Max at 80.6%. It also reports 91.6% on LiveCodeBench for Flash Max and 93.5% for Pro Max. Those figures don't say V4-Flash beats the Pro line. They say the cheap model is close enough to make the expensive model work for its place. Artificial Analysis gives another useful check. Its DeepSeek V4-Flash page lists the model as released in April 2026, with a score of 40 on its Intelligence Index, 118.2 output tokens per second, a one-million-token context window and the same $0.14 and $0.28 API pricing. A separate FundaAI benchmark put DeepSeek V4 Flash at an average 165 seconds per task in its 38-task test, but that wasn't an Artificial Analysis figure. Keep the name attached to the right source. Readers notice when attribution gets lazy. Open-weight models now have enterprise proof The market context is no longer theoretical. TechCrunch reported that Together AI raised $800 million at an $8.3 billion valuation on July 1, with annual bookings above $1.15 billion. The same report said open-source model usage had tripled across the industry over the prior year, citing Together AI's claim based on OpenRouter research. You can argue over the exact category label, open source, open weight, neocloud, you name it. The invoices are less subtle. Companies are paying for cheaper infrastructure that lets them run strong models without sending every task to a closed frontier API. DeepSeek V4-Flash fits that shift neatly because its weights are available for self-hosting under an MIT license, according to the Hugging Face listing. The API is the easy route. It isn't the only route. That difference matters for teams with privacy rules, custom latency needs or enough volume to justify running their own stack. OpenAI and Anthropic still have advantages that don't fit into a token price table. Trust, product maturity, enterprise support and safety workflows matter when you're putting these systems near customer data or internal code. But the buyer's question has changed. You used to ask whether open-weight models were good enough to test. Now you ask where the closed model is still worth the premium. For teams sitting on six-figure AI contracts, the next step is plain: run your own evals against the work you actually pay for. DeepSeek V4-Flash doesn't need to win every benchmark to change your budget. It only needs to be good enough on the tasks you repeat every day. Also read: Kioxia's 31-fold profit surge still wasn't enough to satisfy Wall Street https://startupfortune.com/kioxias-31-fold-profit-surge-still-wasnt-enough-to-satisfy-wall-street/ • Revolut is turning ChatGPT Go into a bank account perk for 75 million customers https://startupfortune.com/revolut-is-turning-chatgpt-go-into-a-bank-account-perk-for-75-million-customers/ • A federal judge says the Trump administration still hasn't proven its case for restricting Anthropic https://startupfortune.com/a-federal-judge-says-the-trump-administration-still-hasnt-proven-its-case-for-restricting-anthropic/