{"slug": "cognition-s-swe-2-achieves-92-8-on-terminal-bench-2-1", "title": "Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1", "summary": "Cognition released SWE-2, a 2.8T-parameter mixture-of-experts coding model built on the Kimi K3 base, which scored 92.8 on Terminal-Bench 2.1 — the highest figure in the published table — and 50.0 on FrontierCode 1.1 Main, one point behind Claude Fable 5.1 (50.9) and 3.3 behind GPT-6 Astra (53.3). Cognition reports SWE-2 runs at a claimed 64% lower cost than Fable 5.1 and a quarter of Astra's cost, but trails both on Terminal-Bench 4.0 with 27.3 versus 55.8 and 57.9. The weights are proprietary with no per-token API pricing, and SWE-2 is available today in Devin Desktop and CLI, with rollout on Devin Web and Fusion; all figures are Cognition's own and pending independent replication.", "body_md": "# SWE-2\n\nMoE premier\n**2.8T total params, 104B active per token (MoE)** - the Kimi K3 base with Cognition’s post-training on top, and the first time Cognition has scaled RL into the multi-trillion-parameter regime. The base had already been RL-heavy for agentic coding; Cognition’s pass added another 5 to 6 points on most benchmarks.\n\n- \n**Serving stack:** MoE inference on NVFP4 and FP8 kernels with quantization-aware training; FP8 carries K, Q, V, and score computations in the MLA layers. A draft model retrained with SpecForge gives 15% longer accept lengths, and a prefill delayer lifts TPM per GPU and tokens/sec per request by 10 to 20% (TTFT takes the hit).\n- \n**Effort levels:** mean steps per run 53 (medium), 80 (high), 98 (max), against 127 for SWE-1.7. Medium posts a higher FrontierCode score than SWE-1.7 with 58% fewer turns and 81% lower average cost, and lands its first real edit at a median of step 18 (SWE-1.7: 48).\n\n**Benchmarks (Cognition self-reported):** FrontierCode 1.1 Main 50.0, DeepSWE 1.1 73.0, Terminal-Bench 2.1 92.8, Terminal-Bench 4.0 27.3. The headline: 50.0 on FrontierCode is one point behind Claude Fable 5.1 (50.9) and 3.3 behind GPT-6 Astra (53.3) - at a claimed 64% lower cost than Fable 5.1 and a quarter of Astra’s. Terminal-Bench 2.1 is the highest number in the published table. The soft spot is Terminal-Bench 4.0, where SWE-2’s 27.3 trails Fable 5.1 (55.8) and GPT-6 Astra (57.9) by a wide margin - long-horizon agentic work is where the gap to the frontier still lives.\n\n**Proprietary weights, no local run.** Cognition has not published SWE-2 weights, so there is nothing to download and no quant ladder to wait for. It is available today in Devin Desktop and CLI, with rollout on Devin Web and Fusion. Cognition publishes no per-token API for SWE-2, so the cost-per-task comparisons (64% cheaper than Fable 5.1 at FrontierCode parity) are the pricing surface, not a $/1M rate card. Every figure here is Cognition’s own number, pending independent replication.\n\n- 2800.0B\n- proprietary\n- 🇺🇸 USA\n- Sep 2026\n\n## What people are building with SWE-2\n\nReal demos from X\n\n## Benchmark scores\n\nVendor-reported - from the developer's own model card / tech report\n\n## Or run it in the cloud\n\n        No per-token API provider pricing tracked for SWE-2 yet.\n        For flagship list prices, see the\n        [calculator](/calculator).\n      \n\n## Inference cost over time\n\nData accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.", "url": "https://wpnews.pro/news/cognition-s-swe-2-achieves-92-8-on-terminal-bench-2-1", "canonical_source": "https://tokenstead.ai/models/swe-2", "published_at": "2026-09-10 16:52:11+00:00", "updated_at": "2026-09-10 17:37:09.754558+00:00", "lang": "en", "topics": ["large-language-models", "ai-agents", "ai-products", "ai-research", "developer-tools"], "entities": ["Cognition", "SWE-2", "Kimi K3", "Devin Desktop", "Devin Web", "Fusion", "Claude Fable 5.1", "GPT-6 Astra"], "alternates": {"html": "https://wpnews.pro/news/cognition-s-swe-2-achieves-92-8-on-terminal-bench-2-1", "markdown": "https://wpnews.pro/news/cognition-s-swe-2-achieves-92-8-on-terminal-bench-2-1.md", "text": "https://wpnews.pro/news/cognition-s-swe-2-achieves-92-8-on-terminal-bench-2-1.txt", "jsonld": "https://wpnews.pro/news/cognition-s-swe-2-achieves-92-8-on-terminal-bench-2-1.jsonld"}}