Just yesterday, OpenAI had cut the prices of its GPT 5.6 Luna model by 80%, but a Chinese lab has now seemingly responded in a big way.
DeepSeek has pushed the official API for DeepSeek-V4-Flash-0731 into public beta, and the numbers attached to it are hard to ignore. The company says it has made a substantial upgrade to the model’s agentic capabilities, with benchmark scores now well ahead of even its own V4-Pro-Preview. The API now natively supports the Responses format and has been adapted for Codex, putting it in more direct competition with how developers are already building around OpenAI’s tooling.
What makes this release notable is what didn’t change. DeepSeek has been explicit that V4-Flash-0731 keeps the same architecture and parameter count as the preview version — this is a training and post-training upgrade, not a new model. The company also clarified that today’s update applies only to the V4-Flash API. The V4-Pro API and the consumer app and web versions remain on their existing builds for now, with DeepSeek saying an official V4-Pro release is coming soon.
DeepSeek-v4-Flash-0731 Benchmarks #
The benchmark jump from preview to this release is steep. On Terminal Bench 2.1, V4-Flash-0731 scores 82.7, up from 61.8 for the preview build and 72.1 for V4-Pro-Preview. That puts it ahead of Z.AI’s GLM-5.2 at 81.0 and within striking distance of Claude Opus 4.8’s 85.0. The pattern repeats across DeepSeek’s other agent evaluations. On NL2Repo, the new model scores 54.2 against 39.4 for the preview. On Cybergym, it jumps to 76.7 from 38.7. DeepSWE goes from 7.3 to 54.4, and Toolathlon-Verified climbs from 49.7 to 70.3. On DeepSeek’s own internal sets, DSBench-FullStack rises to 68.7 and DSBench-Hard to 59.6, both closing in on Opus 4.8’s 71.6 and 71.7 on the same tests.
DeepSeek-v4-Flash-0731 Pricing #
The comparison that will get the most attention outside DeepSeek’s own charts is pricing. V4-Flash-0731 is priced at $0.14 per million input tokens on a cache miss, $0.0028 on a cache hit, and $0.28 per million output tokens, with a concurrency limit of 2,500. Set against a Terminal-Bench 2.1 score of 82.7, that undercuts GPT-5.6 Sol at $5/$30, Claude Fable 5 at $10/$50, and Claude Sonnet 5 at $3/$15 by a wide margin, while trailing Sol on raw score by only about three points and beating both Fable 5 and Sonnet 5 outright.
The release lands one day after OpenAI cut the price of GPT-5.6 Luna by 80%, bringing it down to $0.20 input and $1.20 output per million tokens, a move widely read as a response to pressure from cheaper Chinese models. DeepSeek’s own pricing on V4-Flash-0731 sits well below even that reduced rate, and the timing of the announcement, a single day later, suggests DeepSeek was either already sitting on this update or moved quickly to get ahead of the news cycle.
DeepSeek has previously been candid about trailing the frontier labs by roughly three to six months on raw capability. V4-Flash-0731 doesn’t close that gap entirely, and Opus 4.8 still leads on every benchmark DeepSeek has published. But the margin has narrowed considerably from where V4-Pro-Preview stood just weeks ago, and it has narrowed on a model that costs a fraction of what any of its frontier competitors charge. For teams running high-volume agentic workloads, where thousands of model calls stack up inside a single task, that combination of near-frontier performance and rock-bottom pricing is likely to matter more than which lab holds the top score on any individual benchmark.
Whether V4-Flash-0731 actually holds up to the same scrutiny that hit earlier V4 releases remains to be seen once independent evaluators like Artificial Analysis get their own numbers out. Still, with the official V4-Pro release now teased as imminent, DeepSeek appears to be setting up a second, larger price and performance move before the year is out.