cd /news/ai-agents/fugu-max-and-fugu-ultra-v2-orchestra… · home topics ai-agents article
[ARTICLE · art-126590] src=sakana.ai ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier

Sakana AI released Fugu Max and Fugu Ultra v2, two configurations of its Fugu orchestration architecture that route tasks across a pool of open and specialized models, including the NVIDIA Nemotron family. Sakana AI said Fugu Max achieves the best overall score on six benchmarks, including Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench, and its internal SWEFish benchmark, at $2 per million input tokens and $6 per million output tokens, with output pricing 40-60% lower than comparable frontier models. Fugu Ultra v2 targets peak capability on complex multi-step tasks without relying on the frontier models it orchestrates.

read6 min views1 publishedSep 11, 2026
Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier
Image: Sakana (auto-discovered)

The AI industry has spent a decade racing along a single axis: building bigger, more expensive foundation models. But the frontier that actually matters to real-world tasks is two-dimensional: capability on one axis, cost on the other.

A system that deploys a multi-trillion-parameter model to execute a simple data lookup is not intelligent, but wasteful. The future belongs to systems that know not just how to solve a problem, but which machinery to deploy for the lowest possible cost.

Today, we are pushing orchestration forward along both axes simultaneously. We are releasing Fugu Max, which expands the Pareto Efficiency Frontier by orchestrating our largest pool of open and specialized models to date. And we are releasing Fugu Ultra v2, which pushes peak performance higher than ever before, without the indispensable reliance on the frontier models it orchestrates.

👉 Try Sakana Fugu Max and Fugu Ultra v2 🐡 Schematic illustration of how Fugu Max and Fugu Ultra extend the cost-performance frontier beyond what single models can reach.

Two Axes, One Strategy #

Fugu Max and Fugu Ultra v2 are not separate products. They are the same core orchestration architecture optimized for two distinct missions:

  • Fugu Max asks: What is the best possible output we can deliver at the lowest possible cost?
  • Fugu Ultra v2 asks: What is the absolute highest capability we can achieve on complex, multi-step tasks?

The Fugu Journey 🐡 #

In just a few months, Sakana Fugu has evolved from a beta thesis into an enterprise-grade orchestration engine:

- **April ([Beta](/fugu-beta/))** : Proved multi-agent orchestration works as a unified foundation model.
- **June ([General Availability & Fugu Ultra v1](/fugu-release/))** : Highlighted that an orchestration layer can match closed frontier models on hard benchmarks.
  • July (Fugu-Cyber & Claude Code Interface) : Showed that orchestration can specialize in real-world domain workflows like cybersecurity and coding environments.

  • August (Sakana Chat & NVIDIA Partnership) : Demonstrated orchestration can scale to daily consumer use in Sakana Chat. Started integratingNVIDIA Nemotron open models.

  • September (Today) : Fugu Max and Fugu Ultra v2 Today’s release confirms the core promise behind every milestone: orchestration consistently outperforms isolated models, and a swappable pool of agents guarantees supply chain resilience by design.

Fugu Max: More Models, Less Cost #

Fugu Max expands the pool of models Sakana Fugu can orchestrate, integrating an unprecedented number of open-weights and specialized models, including NVIDIA Nemotron family through our collaboration with NVIDIA.

By dynamically routing tasks to the leanest model capable of solving them, Fugu Max delivers frontier-grade results at a fraction of the token spend.

Among frontier models in a similar price range (input prices per 1M tokens for each model are shown in the first subplot), Fugu Max expands the pareto frontier formed by single models and places itself in a cost-performance efficient position across multiple benchmarks.

Fugu Max sits at a point on the Pareto frontier that single-model providers cannot reach: performance within striking distance of elite models at two to six times lower cost. And because Fugu is an architecture, not a single model, that frontier point can be tuned and extended for any domain.

Concretely,

  • Performance : Fugu Max achieves best overall score on six benchmarks, including Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench, and SWEFish, our internal benchmark reflecting Sakana AI’s own coding challenges and use-cases.
  • Cost : At $2 per million input tokens and $6 per million output tokens, Fugu Max’s output pricing is 40-60% lower than Sonnet 5, GPT 5.6 Terra, and Kimi K3.
  • Efficiency : Fugu Max expands the cost-performance Pareto frontier on seven out of ten benchmarks, delivering performance beyond the existing baseline efficiency envelope.

We see that open models are the fastest-growing and most diverse part of the AI ecosystem, and they become dramatically more useful when orchestrated together rather than used in isolation. Fugu Max is our bet that the Pareto frontier of the future will be built out of many open, specialized models working in concert.

Fugu Ultra v2: The Frontier Keeps Moving #

Pushing cost efficiency does not mean capping maximum capability. For complex multi-step reasoning, autonomous research, and full-stack software development, Fugu Ultra v2 sets our new benchmark for raw output quality.

Where Fugu Ultra v2 separates from the field is on tasks requiring sustained reasoning over complex visual and structured data. It impressively tops SWEFish, demonstrating power in real-world coding challenges and use-cases. On Chartography, which tests visual reasoning and data interpretation, Fugu Ultra v2 scores 48.3, outperforming Opus 5 at 27.3 and Fable 5 at 29.5. On DeepSWE, a benchmark for real-world software engineering, it scores 74.3, outperforming models that cost three to five times more per token.

Peak performance across hard benchmarks. Our flagship Fugu Ultra model continues to deliver a performance that is better or on-par with frontier models. Note: Fugu Ultra v2's training cutoff date is 20260828, Fable 5, Fable 5.1 and GPT-6-Astra are NOT in Fugu-Ultra v2's model pool.

In summary, Fugu Ultra v2

  • Performance : Achieves the best or joint-best score on five of eight benchmarks: GDP.pdf, Chartography, SWEFish, DeepSWE, and Toolathon.
  • Consistency : Places in the top 2 on seven of eight benchmarks, demonstrating strong performance across a broad range of agentic tasks.
  • Frontier : With a focus different from Fugu Max, Fugu Ultra pushes the Pareto frontier in the performance direction, providing a higher-capability option for workloads where quality is the priority.

Crucially, Fugu Ultra v2 achieves these scores without Fable 5, Fable 5.1, or GPT-6-Astra in its agent pool.

Fugu Ultra v2 does not rely on individual proprietary frontier models to deliver frontier output. By orchestrating a swappable pool of open and specialized models, it outperforms closed ecosystems while protecting users from vendor lock-in, API revocations, geopolitical turbulence, and sudden service cutoffs.

Fugu Max expands the frontier to the northwest. Fugu Ultra v2 pushes it upward. Together, they prove that orchestration is not a trade-off between cost and capability. It is the architecture that optimizes both simultaneously.

Both models are available today via our standard OpenAI-compatible API.

If you are already running Fugu, upgrading to Max or Ultra v2 requires a single-line parameter change. No migration. New architecture, same API.

To get started, visit our [product page](/fugu/) or [console site](https://console.sakana.ai).

Orchestration for Everyone #

We believe the most capable AI will never come from a single monolithic model. It will come from intelligent, collective orchestration.

With Fugu Max driving down the cost of intelligence and Fugu Ultra v2 pushing the boundaries of autonomous execution, Sakana Fugu provides the resilient, vendor-agnostic infrastructure required for true AI sovereignty.

── more in #ai-agents 4 stories · sorted by recency
── more on @sakana ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/fugu-max-and-fugu-ul…] indexed:0 read:6min 2026-09-11 ·