cd /news/artificial-intelligence/deepseek-s-v4-flash-undercuts-openai… · home topics artificial-intelligence article
[ARTICLE · art-83655] src=startupfortune.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

DeepSeek's V4-Flash Undercuts OpenAI and Anthropic on Price Again

DeepSeek's V4-Flash API pricing undercuts OpenAI and Anthropic, listed at $0.14 per million cache-miss input tokens and $0.28 per million output tokens, compared to OpenAI's GPT-5.6 Luna at $0.20 input and $1.20 output and Anthropic's Haiku 4.5 at $1 input and $5 output. The model uses a 284 billion parameter mixture-of-experts architecture with 13 billion active parameters and a 1 million token context window, but benchmarks show it trails Kimi K3 significantly, with Terminal-Bench 2.0 scores of 49.1% versus 88.3% and GPQA scores of 71.2% versus 93.5%. DeepSeek's strategy relies on Huawei Ascend 950 chips to lower compute costs, pressuring American labs' premium pricing.

read4 min views3 publishedAug 2, 2026
DeepSeek's V4-Flash Undercuts OpenAI and Anthropic on Price Again
Image: Startupfortune (auto-discovered)

DeepSeek's V4-Flash is brutally cheap, but the stronger story isn't that it beat every higher-priced model. The story is that the price floor for useful AI keeps falling.

DeepSeek doesn't need a perfect benchmark sweep to make OpenAI and Anthropic uncomfortable. Its current API pricing page lists DeepSeek-V4-Flash at $0.14 per million cache-miss input tokens, $0.0028 per million cached input tokens and $0.28 per million output tokens. That part checks out.

The old version of this story went too far. It said DeepSeek pushed V4-Flash-0731 out of preview on July 31 and made its cheapest model beat its own flagship. I couldn't verify either claim. DeepSeek's own changelog says V4-Pro and V4-Flash became available through the API on April 24, 2026, and its model page still identifies V4-Flash as the cheaper, smaller option inside the V4 family. Cheap is not the same as best.

The price comparison is still sharp enough. Axios reported that OpenAI cut GPT-5.6 Luna pricing on July 30 to $0.20 per million input tokens and $1.20 per million output tokens, although OpenAI's own model documentation has also shown Luna at $1 and $6 on its standard pricing pages. Either way, DeepSeek is under it on listed V4-Flash pricing. Against Anthropic's own Haiku 4.5 pricing, $1 per million input tokens and $5 per million output tokens, DeepSeek is about seven times cheaper on input and nearly 18 times cheaper on output.

That is the hook. For developers building coding assistants or document tools that chew through millions of tokens, output pricing isn't a rounding error. It becomes the monthly bill.

DeepSeek's official V4 preview note gives the technical reason it can sell the model this way. V4-Flash uses a 284 billion parameter mixture-of-experts architecture with 13 billion active parameters, supports a 1 million token context window and can run in thinking or non-thinking mode. V4-Pro is much larger, with 1.6 trillion total parameters and 49 billion active parameters, but Flash is the high-volume route. You use it when cost and context length matter more than squeezing out every last benchmark point.

The benchmark story is messier #

The previous draft leaned on an unverified Terminal Bench 2.1 score of 82.7 and treated that as proof that V4-Flash had beaten V4-Pro-Preview. That claim had to go. BenchLM's public page for DeepSeek V4 Flash, verified July 28, gives a more modest picture: an overall score of 57.98, a Terminal-Bench 2.0 score of 49.1 percent and limited benchmark coverage across 22 of the 369 benchmarks it tracks.

Kimi K3 is the harder comparison. BenchLM's DeepSeek V4 Flash versus Kimi K3 page shows Kimi winning Terminal-Bench 2.0 by 88.3 percent to 49.1 percent and GPQA by 93.5 percent to 71.2 percent. That isn't a close call. Still, Kimi's listed API price on the same comparison is $3 per million input tokens and $15 per million output tokens, so DeepSeek's argument is not that Flash wins every test. It is that the work may not need the winner.

That's a real position. If you're routing simple agent steps, summarization passes or repetitive coding support through an API, the best model on the leaderboard may be wasted money. Use the expensive model where it earns the spend. Don't use it as a reflex.

The compute bet behind the discount #

DeepSeek's price strategy sits on a hardware bet. Fortune reported in April that DeepSeek's V4 launch came with close integration with Huawei chips, and DeepSeek expected to cut V4-Pro prices later in the year as Huawei scaled production of its Ascend 950 processor. DeepSeek did not invent cheap inference by writing a cheaper invoice. It is trying to change the cost base underneath the invoice.

That matters for American labs because their premium pricing has to survive two pressures at once: falling Chinese API prices and rising customer discipline over token use. Anthropic can still argue for quality, safety and reliability at the Opus and Sonnet tiers. OpenAI can point to platform depth and distribution. Fine. Those are real advantages.

But when DeepSeek lists V4-Flash output at $0.28 per million tokens, the buyer's question gets much simpler. Why pay more for the easy work?

The answer will not come from a launch post. It will come from invoices, latency charts, failed tasks and the boring production logs developers actually trust. For now, the verified story is narrower than the original article claimed, but it is still important: DeepSeek hasn't proved V4-Flash is better than every higher-priced rival. It has proved that useful AI can keep getting cheaper, and that alone is enough to make the premium labs sweat.

Also read: Orchid's viral anniversary ad backfires into a debate over AI and intimacyNvidia's AI Chip Demand Is Outpacing Supply 12 to 1, Says Dan IvesA Hacker Turned DeepSeek Into an Autonomous Weapon Against 460 Servers

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @deepseek 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/deepseek-s-v4-flash-…] indexed:0 read:4min 2026-08-02 ·