{"slug": "deepseek-v4-flash-beats-its-own-pro-model-on-agent-benchmarks-same-architecture", "title": "DeepSeek V4 Flash Beats Its Own Pro Model on Agent Benchmarks — Same Architecture, Radical Post-Training Gains", "summary": "DeepSeek shipped the public beta of V4 Flash on July 31 with no architecture changes — same 284B MoE, 1M-token context — but a retraining-only round boosted Terminal Bench 25.8 points to 82.7, Toolathlon 18.5 points to 70.3, and delivered first-time scores on Cybergym (76.7), DeepSWE (54.4), and NL2Repo (54.2). Flash now exceeds V4 Pro Preview on several agent tasks at a tenth the cost, marking the first time a 'Flash' tier model has beaten a 'Pro' tier model from the same lab on agent benchmarks.", "body_md": "# DeepSeek V4 Flash Beats Its Own Pro Model on Agent Benchmarks — Same Architecture, Radical Post-Training Gains\n\nDeepSeek shipped the public beta of V4 Flash on July 31 with no architecture changes — same 284B MoE, 1M-token context — but a retraining-only round that boosted Terminal Bench 25.8 points to 82.7, Toolathlon 18.5 points to 70.3, and delivered first-time scores on Cybergym (76.7), DeepSWE (54.4), and NL2Repo (54.2). Flash now exceeds V4 Pro Preview on several agent tasks at a tenth the cost, marking the first time a 'Flash' tier model has beaten a 'Pro' tier model from the same lab on agent benchmarks.\n\n[DeepSeek](/compare/llama-4-vs-deepseek-r1) shipped the public beta of V4 Flash on July 31, and the numbers are strange. Not the architecture — same 284B MoE, same 1M-token context, same price point. What changed was a retraining-only round that turned a decent efficiency model into something that beats DeepSeek's own V4 Pro on agentic tasks.\n\n## The Numbers\n\nTerminal Bench jumped 25.8 points to 82.7. Toolathlon gained 18.5 points, hitting 70.3. Cybergym, DeepSWE, and NL2Repo — three agent benchmarks the preview couldn't even score on — now register 76.7, 54.4, and 54.2 respectively. On internal full-stack dev tests, Flash now exceeds V4 Pro Preview.\n\nThe model costs $0.28 per million input tokens and $1.10 per million output tokens — roughly the same tier as [Gemini 2](/compare/gpt-4o-vs-gemini-2-pro).5 Flash and GPT-5 Mini. But the agent scores suggest something closer to frontier pricing at a fraction of the cost.\n\n## What This Means\n\nPost-training as a differentiator isn't new. But the gap between what V4 Flash preview could do and what the 0731 build can do is unusually wide. It suggests DeepSeek's team found leverage in agent-specific [fine-tuning](/glossary/fine-tuning) that hadn't been fully explored — and that the next generation of benchmarks will measure agent capability, not raw [reasoning](/glossary/reasoning).\n\nThe model stayed in preview for months while the post-training team worked. The result is a model that writes code, uses tools, and completes multi-step tasks better than the more expensive, larger sibling it was supposed to sit beneath.\n\n## The Competitive Angle\n\n[Anthropic](/glossary/anthropic)'s Opus 4 leads on high-end reasoning. OpenAI's GPT-5 Pro owns latency-sensitive enterprise workloads. But DeepSeek is carving a specific lane: agent capability without the agent tax. If Flash can do what Pro does on tool use at a tenth the cost, enterprise [inference](/glossary/inference) budgets shift.\n\nDeepSeek's consumer app and web endpoints are unaffected. The API change is transparent — same model name, same call signature, dramatically different behavior on anything involving tools, code, or multi-turn task execution.\n\n## The Open Question\n\nWhether a post-training-only improvement is sustainable. If the gains came from targeted [synthetic data](/glossary/synthetic-data), competitors replicate it in weeks. If they came from a fundamentally better RL pipeline, DeepSeek has an edge that compounds.\n\nEither way, July 31 marks the first time a \"Flash\" model legitimately beat a \"Pro\" model from the same lab on agent benchmarks. That won't be the last.\n\nGet AI news in your inbox\n\nDaily digest of what matters in AI.\n\n## Key Terms Explained\n\n[Anthropic](/glossary/anthropic)\n\nAn AI safety company founded in 2021 by former OpenAI researchers, including Dario and Daniela Amodei.\n\n[Fine-Tuning](/glossary/fine-tuning)\n\nThe process of taking a pre-trained model and continuing to train it on a smaller, specific dataset to adapt it for a particular task or domain.\n\n[Gemini](/glossary/gemini)\n\nGoogle's flagship multimodal AI model family, developed by Google DeepMind.\n\n[GPT](/glossary/gpt)\n\nGenerative Pre-trained Transformer.", "url": "https://wpnews.pro/news/deepseek-v4-flash-beats-its-own-pro-model-on-agent-benchmarks-same-architecture", "canonical_source": "https://www.machinebrief.com/news/deepseek-v4-flash-beats-pro-model-agent-benchmarks-terminal-bench-toolathlon-july-2026", "published_at": "2026-07-31 13:07:14+00:00", "updated_at": "2026-07-31 13:33:58.373795+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-products", "ai-agents"], "entities": ["DeepSeek", "V4 Flash", "V4 Pro Preview", "Terminal Bench", "Toolathlon", "Cybergym", "DeepSWE", "NL2Repo"], "alternates": {"html": "https://wpnews.pro/news/deepseek-v4-flash-beats-its-own-pro-model-on-agent-benchmarks-same-architecture", "markdown": "https://wpnews.pro/news/deepseek-v4-flash-beats-its-own-pro-model-on-agent-benchmarks-same-architecture.md", "text": "https://wpnews.pro/news/deepseek-v4-flash-beats-its-own-pro-model-on-agent-benchmarks-same-architecture.txt", "jsonld": "https://wpnews.pro/news/deepseek-v4-flash-beats-its-own-pro-model-on-agent-benchmarks-same-architecture.jsonld"}}