# Devin Fusion: the first multi-model coding agent on the Pareto frontier

> Source: <https://twitter.com/ArtificialAnlys/status/2098504939781906684>
> Published: 2026-09-22 12:39:16+00:00

Artificial Analysis on X: "Devin Fusion performs well for cost efficiency and performance, and currently sits on the Pareto frontier for Coding Agent Index score vs. Cost per Task in both configurations we tested
Devin Fusion CLI with Claude Fable 5.1 (xhigh) + SWE-2 (medium) scores 61.7 on the Ar… / X

Artificial Analysis on X: "Devin Fusion performs well for cost efficiency and performance, and currently sits on the Pareto frontier for Coding Agent Index score vs. Cost per Task in both configurations we tested
Devin Fusion CLI with Claude Fable 5.1 (xhigh) + SWE-2 (medium) scores 61.7 on the Artificial Analysis Coding Agent Index v1.5. This is almost tied with Claude Fable 5.1 (max, with fallback) in Claude Code at 62.2 despite the lower effort, and costs 36% less at $7.9 per task for the Fusion configuration vs. $12.4 for Claude Code. Speed is also essentially flat, with time per task of 35.8 vs. 34.8 minutes.
This pattern holds across the underlying evaluations: Fusion scores 63.1 vs. 64.3 on DeepSWE 1.1, 65.9 vs. 64.8 on SWE-Atlas QnA, and 56.1 vs. 57.6 on Terminal-Bench 4.0."

Devin Fusion performs well for cost efficiency and performance, and currently sits on the Pareto frontier for Coding Agent Index score vs. Cost per Task in both configurations we tested
Devin Fusion CLI with Claude Fable 5.1 (xhigh) + SWE-2 (medium) scores 61.7 on the Artificial Analysis Coding Agent Index v1.5. This is almost tied with Claude Fable 5.1 (max, with fallback) in Claude Code at 62.2 despite the lower effort, and costs 36% less at $7.9 per task for the Fusion configuration vs. $12.4 for Claude Code. Speed is also essentially flat, with time per task of 35.8 vs. 34.8 minutes.
This pattern holds across the underlying evaluations: Fusion scores 63.1 vs. 64.3 on DeepSWE 1.1, 65.9 vs. 64.8 on SWE-Atlas QnA, and 56.1 vs. 57.6 on Terminal-Bench 4.0.

We independently benchmarked Devin Fusion for its release today - this is the first time a multi-model coding agent has been included on the Artificial Analysis Coding Agent Index, and it effectively retains Claude Fable 5.1 and GPT-6 Astra performance while reducing costs
DevinShow more

Devin Fusion performs well for cost efficiency and performance, and currently sits on the Pareto frontier for Coding Agent Index score vs. Cost per Task in both configurations we tested
Devin Fusion CLI with Claude Fable 5.1 (xhigh) + SWE-2 (medium) scores 61.7 on the Artificial Analysis Coding Agent Index v1.5. This is almost tied with Claude Fable 5.1 (max, with fallback) in Claude Code at 62.2 despite the lower effort, and costs 36% less at $7.9 per task for the Fusion configuration vs. $12.4 for Claude Code. Speed is also essentially flat, with time per task of 35.8 vs. 34.8 minutes.
This pattern holds across the underlying evaluations: Fusion scores 63.1 vs. 64.3 on DeepSWE 1.1, 65.9 vs. 64.8 on SWE-Atlas QnA, and 56.1 vs. 57.6 on Terminal-Bench 4.0.

Astra and SWE together lands perfectly in the green zone. BTW Fusion works in both CLI and Devin Desktop (if you are more of a GUI guy like myself)
Fusion is a great concept and it's not a simple model router - encourage to read cognition.com/blog/devin-fus…
This is the futureShow more
