{"slug": "deepseek-v4-flash-0731-82-7-on-terminal-bench-2-1-with-a-public-harness", "title": "DeepSeek V4 Flash 0731: 82.7% on Terminal-Bench 2.1 with a public harness", "summary": "DeepSeek's V4 Flash 0731 model scored 82.7% (±1.79 SE) on Terminal-Bench 2.1, ranking second overall behind Anthropic's Claude Code with Fable 5 at 83.8% (±1.16 SE), according to Ante's public benchmark harness. The model's run cost $68.41 and took 38.9 minutes, with results verified using consistent parameters.", "body_md": "Terminal-Bench 2.1\n\nCompare Ante runs across models on the same Terminal-Bench 2.1 task set, using consistent parameters and verified benchmark results.\n\n| # | Model | Same-model | Agent | Source | |||||\n|---|---|---|---|---|---|---|---|---|---|\n| 1 | DeepSeek V4 Flash 0731max | 82.7%±1.79 SE | $68.41 | 38.9 min | #1 same-model | Ante0.preview.71 |\n|\n\nTerminal-Bench Reference\n\nFor how different models perform on TB 2.1, see [Vals AI's Terminal-Bench 2.1 benchmark](https://www.vals.ai/benchmarks/terminal-bench-2-1).\n\n| # | Agent | Model | Accuracy | Run date | |\n|---|---|---|---|---|---|\n| 1 | Claude CodeAnthropic | Fable 5xhigh | 83.8%±1.16 SE | Jun 7, 2026 | |\n| 2 | CodexOpenAI | GPT-5.5xhigh | 83.2%±1.13 SE | May 1, 2026 | |\n| 3 | Terminus 2Terminal-Bench | Fable 5high | 80.5%±1.16 SE | Jun 5, 2026 | |\n| 4 | Cursor CLICursor | Grok 4.5high | 79.3%±1.46 SE | Jul 9, 2026 | |\n| 5 | Claude CodeAnthropic | Opus 4.8high | 78.9%±1.31 SE | Jul 9, 2026 | |\n| 6 | CodexOpenAI | GPT-5.6 Terramax | 78.4%±1.25 SE | Jul 11, 2026 | |\n| 7 | Terminus 2Terminal-Bench | GPT-5.5xhigh | 78.0%±1.22 SE | May 1, 2026 | |\n| 8 | mini-SWE-agentPrinceton | Muse Spark 1.1xhigh | 76.2%±1.23 SE | Jul 9, 2026 | |\n| 9 | CodexOpenAI | GPT-5.6 Lunamax | 75.7%±1.32 SE | Jul 11, 2026 | |\n| 10 | Claude CodeAnthropic | Sonnet 5high | 74.6%±1.64 SE | Jul 9, 2026 | |\n| 11 | Terminus 2Terminal-Bench | Gemini 3 Prohigh | 73.9%±1.29 SE | May 1, 2026 | |\n| 12 | Claude CodeAnthropic | Opus 4.7max | 68.9%±1.41 SE | May 1, 2026 | |\n| 13 | Terminus 2Terminal-Bench | Opus 4.7max | 66.1%±1.37 SE | May 1, 2026 | |\n| 14 | Gemini CLIGoogle | Gemini 3 Prohigh | 65.8%±1.38 SE | May 1, 2026 | |\n| 15 | Gemini CLIGoogle | Gemini 3.1 Prohigh | 65.8%±1.67 SE | May 5, 2026 | |\n| 16 | Terminus 2Terminal-Bench | Gemini 3.1 Prohigh | 65.6%±1.65 SE | May 5, 2026 | |\n| 17 | Claude CodeAnthropic | GLM-5.1max | 58.6%±1.24 SE | May 1, 2026 |", "url": "https://wpnews.pro/news/deepseek-v4-flash-0731-82-7-on-terminal-bench-2-1-with-a-public-harness", "canonical_source": "https://antigma.ai/eval", "published_at": "2026-08-09 08:44:29+00:00", "updated_at": "2026-08-09 09:01:20.496112+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "ai-products"], "entities": ["DeepSeek", "Ante", "Terminal-Bench", "Anthropic", "Claude Code", "Fable 5", "Vals AI"], "alternates": {"html": "https://wpnews.pro/news/deepseek-v4-flash-0731-82-7-on-terminal-bench-2-1-with-a-public-harness", "markdown": "https://wpnews.pro/news/deepseek-v4-flash-0731-82-7-on-terminal-bench-2-1-with-a-public-harness.md", "text": "https://wpnews.pro/news/deepseek-v4-flash-0731-82-7-on-terminal-bench-2-1-with-a-public-harness.txt", "jsonld": "https://wpnews.pro/news/deepseek-v4-flash-0731-82-7-on-terminal-bench-2-1-with-a-public-harness.jsonld"}}