DeepSeek's New V4-Flash-0731 Undercuts OpenAI's GPT-5.6 on Price DeepSeek released V4-Flash-0731 on July 31, a retrained version of its budget model that outperforms its own V4-Pro-Preview on all nine agent and coding benchmarks, scoring 82.7 on Terminal-Bench 2.1 versus 72.1 for V4-Pro-Preview. The model costs $0.14 per million input tokens on a cache miss, undercutting OpenAI's GPT-5.6 Luna by about 30% on input cost, and lands about one point behind on benchmarks while running at roughly 60% lower cost per task. The benchmarks are DeepSeek's own, not independently verified, but Artificial Analysis's Intelligence Index confirms a gain. DeepSeek just made its cheapest model beat its own flagship, and it undercuts OpenAI's newest pricing by roughly 30 percent. DeepSeek pushed a new build of its budget model live on July 31, moving DeepSeek-V4-Flash out of preview and into an official public beta under the designation V4-Flash-0731. The company didn't just flip a switch. It retrained the model, and the retrained version now beats DeepSeek's own pricier V4-Pro-Preview on all nine agent and coding benchmarks the company published, according to MarkTechPost. On Terminal-Bench 2.1, a test of how well a model handles real terminal and coding tasks, V4-Flash-0731 scores 82.7, up from 61.8 for the earlier preview build and ahead of V4-Pro-Preview's 72.1. That's not a small jump. Digital Applied and MarkTechPost both reported a 645% gain on DeepSWE, one of the nine agent benchmarks in the release, and Artificial Analysis put V4-Flash-0731 at 50 on its Intelligence Index, ten points above the previous V4-Flash build. The model's architecture hasn't changed. It's still a 284-billion-parameter mixture-of-experts system with 13 billion active parameters per token, the same footprint as the April preview. DeepSeek simply re-post-trained it, and the gains came entirely from that process. Here's why the timing matters. The 0731 release landed one day after OpenAI cut the price of GPT-5.6 Luna by 80%, according to The Decoder, bringing it down to $0.20 per million input tokens and $1.20 per million output tokens. V4-Flash-0731 costs $0.14 per million input tokens on a cache miss, just $0.0028 on a cache hit, and $0.28 per million output tokens. That undercuts OpenAI's freshly discounted Luna tier by about 30% on input cost, and it beats Anthropic's Claude Haiku 4.5 by a factor of seven. On raw capability, The Decoder found V4-Flash-0731 lands about one point behind GPT-5.6 Luna on benchmark scores while running at roughly 60% lower cost per task. Frankly, that's DeepSeek's whole strategy in one line: get close enough on capability that price is the only argument left. The new API also natively supports OpenAI's Responses format and has been adapted to work with Codex, MarkTechPost noted, which puts DeepSeek in more direct competition for developers who already built their workflows around OpenAI's tooling. That's a deliberate move to make switching cheap, not just the tokens themselves. The catch: these are DeepSeek's own numbers One caveat worth stating plainly. Every benchmark above, Terminal-Bench 2.1, DeepSWE, the nine-benchmark sweep, comes from DeepSeek's own testing harness. As of July 31, none of it had been independently reproduced by a third party. Artificial Analysis's Intelligence Index score is the one external data point in the mix, and it still shows a real gain over the previous V4-Flash. That's not nothing. But it's also not the same as a neutral lab running the tests itself. This is DeepSeek's second consequential move on the V4-Flash line in a matter of months. V4 and V4-Pro first shipped as previews on April 24, and V4-Flash followed as the lightweight option built for high-volume, latency-sensitive workloads. Friday's release takes that model out of preview status entirely and puts a version in front of paying developers that, on DeepSeek's own tests, now outperforms the pricier model sitting above it in the lineup. Competitors aren't sitting still either. Z.AI's GLM-5.2 sits close behind V4-Flash-0731 on Terminal-Bench at 81.0, and the DeepSeek model trails Claude Opus 4.8's 85.0 by just over two points. Moonshot AI, meanwhile, is reportedly scaling up with 20,000 new Nvidia GPUs, according to wccftech. None of that changes the number that will actually move developer traffic: a 1-million-token context window at 14 cents per million input tokens is a hard figure for OpenAI and Anthropic to answer without cutting their own margins further. OpenAI cut prices Thursday. DeepSeek answered Friday. Also read: OpenAI's and Anthropic's AI Agents Escaped Testing and Hacked Real Firms https://startupfortune.com/openais-and-anthropics-ai-agents-escaped-testing-and-hacked-real-firms/ • Reddit's Stock Sinks After Earnings as Google's AI Overviews Eat Its Traffic https://startupfortune.com/reddits-stock-sinks-after-earnings-as-googles-ai-overviews-eat-its-traffic/ • Google pulled its Earth AI image tool in under 24 hours after users faked disasters and military strikes https://startupfortune.com/google-pulled-its-earth-ai-image-tool-in-under-24-hours-after-users-faked-disasters-and-military-strikes/