Photo: Tima Miroshnichenko / Pexels
The dual-model architecture delivers 46% cost savings and matches top-tier coding benchmarks by splitting the thinking from the doing
Cognition has rolled out Devin Fusion, a coding agent architecture that pairs two AI models together: one does the strategic thinking, the other handles the grunt work.
Fusion works by assigning a frontier-class “lead” model to handle planning and reasoning while a cheaper “sidekick” model tackles execution. The two run in parallel with separate persistent contexts, which means the planner can keep scheming while the executor is busy writing code.
The numbers behind the pairing #
Cognition reports that Fusion delivers up to 39% better efficiency compared to traditional single-model setups. Under specific benchmarking conditions, that translates to a 46% cost reduction.
On the Artificial Analysis Coding Agent Index v1.5, Fusion’s performance matches Claude Fable 5.1 and outperforms GPT-6 Astra.
The recommended configuration pairs Fable 5.1 as the lead model with SWE-2 as the sidekick, though the system also supports GPT-6 Astra pairings.
Cognition says 88% of merged pull requests from its own engineering team were successfully processed through Fusion’s automated routing.
From preview to desktop #
Devin Fusion was first announced on June 29, 2026, as a preview feature. By September 11, 2026, Cognition expanded deployment to the Devin Desktop and CLI interfaces, making the dual-model architecture available across its main product surfaces.
Fusion avoids cache misses by maintaining separate persistent contexts for each model and dynamically routing tasks mid-session.
Cognition’s bigger picture #
The company’s revenue run-rate grew from $73M to over $500M following its merger with Windsurf. Its valuation climbed from $26B in May 2026 to $48B shortly after.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our