DeepSeek has quietly rolled out an official build of V4 Pro, and Artificial Analysis has now put out its independent evaluation of the model. DeepSeek V4 Pro 0813 (max) lands at 53 on the Artificial Analysis Intelligence Index, a composite score built from ten separate benchmarks spanning reasoning, coding, agentic tool use and knowledge. That places it just one point ahead of DeepSeek V4 Flash 0731, the company’s cheaper, faster model that had briefly matched or beaten Pro on the same index. It’s a modest gain on paper, but the context around it, cost, speed, and where it sits on Artificial Analysis’s Pareto frontier, is what makes the release worth a closer look.
For a bit of background, DeepSeek V4 Pro was originally released back in April alongside V4 Flash, debuting at a score of 52 and taking the number two spot among open-weights reasoning models at the time. That release used a 1.6 trillion parameter Mixture-of-Experts architecture with 49 billion active parameters, more than double the size of V3. Since then, DeepSeek pushed out V4 Flash 0731 as an upgrade that closed most of the gap with Pro, and now this 0813 build brings Pro back out in front, if only barely. Where DeepSeek V4 Pro Sits On The Leaderboard
At 53 on the Intelligence Index, DeepSeek V4 Pro (max) trails the current frontier of closed models by a significant margin. Claude Opus 5 (max) leads the field at 63, followed closely by Claude Fable 5 at 62, GPT-5.6 Sol (max) and Grok 4.6 (high) tied at 61, and Kimi K3 (max) at 60. Muse Spark 1.2 (xhigh) sits at 57. DeepSeek V4 Pro then shares its 53 with Z.AI’s GLM-5.2 (max), just ahead of GPT-5.6 Luna (max) and Gemini 3.6 Flash, both at 52, and well clear of Nemotron 3 Ultra at 38.
That gap to the top of the board is real, roughly ten points separating DeepSeek from Claude Opus 5. But the Intelligence Index rewards raw capability regardless of cost, and cost is exactly where DeepSeek’s newest model tries to make its case.
The individual benchmark breakdowns tell a more nuanced story than the single composite number. On GDPval-AA v2, Artificial Analysis’s test for agentic real-world work tasks, DeepSeek V4 Pro scores 55%, ahead of GPT-5.6 Luna’s 53% and Gemini 3.6 Flash’s 50%, though still behind Claude Opus 5’s 67% and Grok 4.6’s 62%. On GPQA Diamond, a graduate-level science reasoning benchmark, DeepSeek V4 Pro actually posts a 93%, tying Claude Opus 5 and Kimi K3 (low), and trailing the category leader Grok 4.6 (high) at 95% by only two points. That’s a strong result for a model sitting mid-table on the overall index.
Coding and terminal-use benchmarks are less flattering. On Terminal-Bench v2.1, which tests agentic terminal workflows, DeepSeek V4 Pro scores 79%, a full ten points behind Claude Opus 5’s 89%. On SciCode, a coding benchmark, it comes in at 49%, the lowest score among the models Artificial Analysis highlighted in this comparison, trailing Claude Fable 5’s 60% and even Nemotron 3 Ultra’s 40%. Agentic tool use shows a similar pattern: on τ³-Banking, DeepSeek V4 Pro scores 40%, ahead of Claude Fable 5’s 39% but behind Grok 4.6’s 51%.
Knowledge and reasoning benchmarks land somewhere in the middle. On Humanity’s Last Exam, DeepSeek V4 Pro scores 39%, tied with MiniMax-M3 and just behind GPT-5.6 Luna. On CritPt, a physics reasoning benchmark, it manages only 18%, well behind GPT-5.6 Sol’s 32% and Claude Opus 5’s 29%. On AA-Omniscience Accuracy, a knowledge benchmark, it scores 49%, roughly in the middle of the pack, behind Claude Fable 5’s 65% but ahead of Grok 4.6’s 48%.
On the more business-facing evaluations, DeepSeek V4 Pro doesn’t appear among the leaders. Artificial Analysis’s AA-Briefcase, which measures agentic knowledge work on an Elo scale, has Claude Opus 5 out front at 1715 Elo, with DeepSeek’s older V4 Pro 0731 build trailing at 1286, well behind Kimi K3 (max) at 1541 and Muse Spark 1.2 at 1358.
Pricing Is Where The Model Actually Wins
DeepSeek’s pitch has never really been about topping the intelligence charts, it’s been about the price attached to whatever score it does put up, and that pattern holds here. DeepSeek V4 Pro is priced at $0.435 per million input tokens and $0.87 per million output tokens, with cache-hit pricing dropping input costs by 99% to $0.004 per million tokens. On Artificial Analysis’s weighted cost-per-task metric, that works out to roughly $0.06 per Intelligence Index task, the second-cheapest figure on the entire chart, behind only GPT-5.6 Luna’s $0.05.
Compare that to what the frontier costs elsewhere. Claude Opus 5 (max) runs $2.34 per task, Claude Fable 5 comes in at $3.14, and GPT-5.6 Sol (max) sits at $1.23. Even Grok 4.6 (high) and Kimi K3 (max), both priced closer to the middle of the pack, land at $0.84 apiece, fourteen times what DeepSeek is charging for a model that trails Opus 5 by only ten points on the index. On Artificial Analysis’s own product page, DeepSeek V4 Pro ranks #2 out of 104 models on intelligence but only #43 on price, and #8 on cache-hit pricing specifically, underlining just how aggressively the caching discount is being used to undercut the field.
Speed is a secondary but related advantage. DeepSeek V4 Pro outputs 83 tokens per second, ranked #18 out of 104 models on Artificial Analysis’s speed metric, well ahead of Claude Opus 5’s 52 tokens per second and Kimi K3’s 41, though behind Gemini 3.6 Flash’s 229 and GPT-5.6 Luna’s 150.
Where DeepSeek V4 Pro Sits On The Pareto Frontier
Artificial Analysis plots every model on a chart of Intelligence Index score against cost per task, and the models that sit on the “Pareto frontier,” the dotted line connecting whichever model delivers the most intelligence for a given price, are the ones worth paying attention to regardless of where they land on the raw leaderboard. DeepSeek V4 Pro 0813 (max) sits directly on that line, at the corner where the frontier flattens out around the $0.06 mark before climbing steeply toward the far more expensive frontier models.
That positioning matters more than the single-point gap over Flash. DeepSeek V4 Flash 0731 (max) sits just behind Pro on the same frontier at a lower cost, meaning DeepSeek currently holds two of the most cost-efficient points on the entire chart. Everything above that flattened section, Muse Spark, GLM-5.2, GPT-5.6 Luna, Kimi K3, Grok 4.6, GPT-5.6 Sol, and eventually Claude Opus 5 and Fable 5, costs progressively more for progressively smaller intelligence gains. Getting from DeepSeek’s 53 up to Claude Opus 5’s 63 means paying roughly 39 times more per task for a ten-point jump on the index.
This is broadly consistent with the pattern DeepSeek has followed since its earlier V4 releases, and the company has repeated the trick with each iteration. When V4 Flash 0731 launched in July, it scored 50 on the index while sharing identical pricing with its predecessor, which meant the entire gain showed up as a near-vertical jump on the same cost chart rather than a shift sideways into a more expensive tier. The 0813 build of Pro appears to be doing something similar, a small further gain in intelligence without a corresponding jump in price.
The Bigger Picture
None of this puts DeepSeek V4 Pro in contention for the top of the Intelligence Index, that spot currently belongs to Claude Opus 5, with GPT-5.6 Sol, Grok 4.6 and Kimi K3 rounding out the top tier. What the 0813 update does is reinforce DeepSeek’s position as the cheapest way to get reasonably close to frontier performance, particularly on reasoning-heavy tasks like GPQA Diamond where it’s nearly tied with the leaders despite the ten-point gap on the composite score.
For businesses weighing intelligence against inference cost at scale, that one-point gain over Flash is less the headline than the fact that DeepSeek now has two models, Pro and Flash, sitting almost on top of each other on the Pareto frontier, both cheaper than nearly everything else on the chart. Whether that’s enough to pull developers away from the frontier labs will likely come down to how much the remaining ten-point intelligence gap actually matters for a given use case, and for high-volume, cost-sensitive workloads, DeepSeek is betting it won’t.