{"slug": "cognition-ships-swe-2-a-cheaper-coding-model-for-devin", "title": "Cognition ships SWE-2, a cheaper coding model for Devin", "summary": "Cognition released SWE-2 on September 10, a coding model it says approaches Fable 5.1's performance at 64% lower mean rollout cost, according to the company's own FrontierCode benchmark. SWE-2 scored 50.0% on the Main set versus Fable 5.1 at 50.9%, GPT-6 Astra at 53.3%, Grok 4.6 at 48.0% and SWE-1.7 at 42.0%, and is available in Devin Desktop and the Devin command-line interface with rollout underway across Devin Web and Fusion. Cognition says it post-trained SWE-2 from Kimi K3, adding five to six percentage points on several coding evaluations, though the claimed lineage and the cost figures come from Cognition's own benchmark and have not been independently confirmed.", "body_md": "# Cognition ships SWE-2, a cheaper coding model for Devin\n\n**Cognition says SWE-2 approaches Fable 5.1's coding performance at 36% of its rollout cost, based on the company's own benchmark.**\n\n        By [RuntimeWire Staff](/author/runtimewire-staff)\n        · Published \n\nPrimary source: [Cognition](https://cognition.com/blog/swe-2)\n\n## Why it matters\n\nSWE-2 tests whether Cognition can turn post-training and tighter agent behavior into lower costs for Devin users. The company's 64% cost claim comes from its own benchmark, making independent evaluation the next meaningful test.\n\nScott Wu's [Cognition](https://cognition.ai/?ref=runtimewire), which operates Devin, released [SWE-2](https://cognition.com/blog/swe-2?ref=runtimewire) on September 10. The model is available in [Devin Desktop](https://devin.ai/desktop?ref=runtimewire) and its [command-line interface](https://devin.ai/cli?ref=runtimewire), with a rollout underway across [Devin Web](https://app.devin.ai/?ref=runtimewire) and [Fusion](https://cognition.com/blog/devin-fusion?ref=runtimewire). Cognition says SWE-2 approaches the coding performance of Fable 5.1 and GPT-6 Astra at substantially lower rollout costs.\n\nThe release turns Wu's argument about software agents into a model-training strategy. He has steered Devin toward engineering work inside large codebases, where completing the task matters more than producing an impressive snippet. SWE-2 is designed to reach the first useful edit faster, spend fewer steps exploring a repository and reserve extended reasoning for harder work.\n\nWu arrived at that thesis after leaving Harvard, working as a founding engineer at Scale AI and co-founding the professional-networking service Lunchclub. He won three International Olympiad in Informatics gold medals and placed first overall in 2014. Cognition co-founder Steven Hao is also an IOI gold medalist, giving the founding team a concentrated group trained to treat software work as a sequence of decisions that can be measured and optimized.\n\n### Cognition post-trains instead of starting over\n\nIn its SWE-2 announcement, Cognition says it post-trained the model from [Kimi K3](https://arxiv.org/abs/2607.24653?ref=runtimewire), though that claimed lineage has not been independently confirmed.\n\nCognition says its post-training added between five and six percentage points on several coding evaluations. The central training change lets Cognition optimize medium, high and maximum reasoning-effort settings in one reinforcement-learning run. Each setting receives a cost penalty calibrated to the local slope of the base model's cost-performance curve, rewarding improvements that raise solve rates without allowing the model to spend freely on longer trajectories.\n\nCognition said it tripled the number of reinforcement-learning environments used for SWE-2, added instruction-following requirements and used earlier SWE-2 checkpoints to identify weaknesses in its automated verifiers. Cognition also trained an online draft model for speculative decoding and used lower-precision NVFP4 and FP8 kernels to hold memory use down while working with a base model almost three times the size of SWE-1.7.\n\n### The benchmark lead comes with a house advantage\n\nOn Cognition's [FrontierCode leaderboard](https://cognition.com/frontiercode?ref=runtimewire), SWE-2 scored 50.0% on the Main set. Cognition reports Fable 5.1 at 50.9%, GPT-6 Astra at 53.3%, [Grok 4.6](/models/azure/grok-4.6) at 48.0% and SWE-1.7 at 42.0%. Cognition says SWE-2's mean rollout cost was 64% lower than Fable 5.1's and roughly one-quarter of GPT-6 Astra's.\n\nThe comparison is best read as Cognition's internal evidence for how SWE-2 behaves inside coding-agent workflows. Cognition created FrontierCode, runs the leaderboard and reports the best score across reasoning-effort settings. Coding models are not mapped one-to-one to agent harnesses: [Moonshot's Kimi K3 documentation](https://github.com/MoonshotAI/Kimi-K3/blob/main/README.md?ref=runtimewire) describes evaluations conducted through Kimi Code and Claude Code, while other models may also run through multiple harnesses.\n\nCognition revised [FrontierCode to version 1.1](https://cognition.com/blog/frontier-code-1.1?ref=runtimewire) in July after auditing more than 1,000 grading criteria, relaxing 75 and changing how internet use is judged. Cognition also retired the Diamond subset. Those changes may improve the evaluation, though they reinforce the need for results outside Cognition's own framework before treating small differences among frontier models as settled rankings.\n\nSWE-2's reported results are uneven enough to make that caution concrete. Cognition reports 73.0% on DeepSWE 1.1, close to GPT-6 Astra's 74.1%, and 92.8% on Terminal-Bench 2.1, ahead of every comparison model listed in the announcement. On Terminal-Bench 4, SWE-2 scored 27.3%, far behind Fable 5.1 at 55.8% and GPT-6 Astra at 57.9%. SWE-2 looks competitive on some agentic coding workloads rather than uniformly equivalent to the leading proprietary models.\n\nThe more practical result may be the reduction in meandering. On the 100-task FrontierCode Main set, Cognition says SWE-2 medium averaged 53 steps per run, down from 127 for SWE-1.7. Its first substantive code edit arrived after a median of 18 steps, compared with 48 for SWE-1.7. Cognition calculates that SWE-2 medium used 58% fewer turns and cost 81% less on average while scoring higher.\n\nFor engineering managers, that behavior can matter as much as another benchmark point. Agents that repeatedly inspect the same files or produce long plans for simple changes consume compute, delay feedback and make their work harder to supervise. Cognition is training SWE-2 to act earlier at medium effort while preserving longer planning and verification at higher settings.\n\n### The release follows a $48B valuation\n\nOn September 8, Cognition [said it had raised more than $2 billion at a $48 billion valuation](/article/cognition-series-e-48-billion-valuation-scott-wu), with Andreessen Horowitz and Accel leading and Founders Fund, General Catalyst and Avenir participating. \n\nIn the [Series E announcement](https://cognition.com/blog/series-e?ref=runtimewire), Cognition described itself as an independent agent lab that can select and combine models instead of binding customers to one provider. Its SWE-2 post describes Cognition applying its own reinforcement-learning system to an outside base model and distributing the result through Devin, where Cognition controls the agent interface and tools.\n\nThe 64% cost advantage remains specific to Cognition's benchmark methodology and mean spend per rollout. Devin customers will judge the economics through completed work and accepted code changes. A model that reaches useful edits sooner could improve those economics, but the reported gains still need validation outside Cognition's evaluation framework.", "url": "https://wpnews.pro/news/cognition-ships-swe-2-a-cheaper-coding-model-for-devin", "canonical_source": "https://runtimewire.com/article/cognition-swe-2-coding-model-scott-wu", "published_at": "2026-09-10 16:44:38+00:00", "updated_at": "2026-09-10 16:51:29.085878+00:00", "lang": "en", "topics": ["ai-products", "ai-tools", "ai-agents", "large-language-models", "ai-research"], "entities": ["Cognition", "SWE-2", "Devin", "Scott Wu", "Steven Hao", "Fable 5.1", "GPT-6 Astra", "Kimi K3"], "alternates": {"html": "https://wpnews.pro/news/cognition-ships-swe-2-a-cheaper-coding-model-for-devin", "markdown": "https://wpnews.pro/news/cognition-ships-swe-2-a-cheaper-coding-model-for-devin.md", "text": "https://wpnews.pro/news/cognition-ships-swe-2-a-cheaper-coding-model-for-devin.txt", "jsonld": "https://wpnews.pro/news/cognition-ships-swe-2-a-cheaper-coding-model-for-devin.jsonld"}}