cd /news/artificial-intelligence/recursive-self-improvement-through-c… · home › topics › artificial-intelligence › article
[ARTICLE · art-149322] src=pub.sakana.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Recursive self-improvement through collective intelligence

Sakana AI researchers introduced Multi-Agent Self-Supervision (MASS), a method that uses a single language model as executor, workflow optimizer, and evaluator to generate a self-supervision signal for recursive self-improvement, distilling a multi-agent team's collective intelligence back into the shared model. MASS improved score per output token over the base model on non-verifiable research benchmarks, with progress measured by separate judgments from GPT-5.5 and Claude Opus 4.8 that were withheld from the training loop. The work was done during an internship at Sakana AI.

read12 min views1 publishedOct 11, 2026
Recursive self-improvement through collective intelligence
Image: source

To advance beyond humanity's accumulated knowledge, AI itself (the optimizee) may also be its best available optimizer and evaluator. MASS repeatedly optimizes multi-agent workflows, then distills collective intelligence from multi-agent systems and uses it as a self-supervision signal for RSI. It improves score per output token over the base model on research benchmarks. We hope MASS contributes to AI’s Odyssean moment: continual self-learning in a setting where the optimizee is the evaluator, and the optimizer is the optimizee.

Estimated reading time: 15 min

† Work done during an internship at Sakana AI.

We want frontier AI models to take us beyond the body of knowledge humanity has built over generations. At that frontier, even human experts may struggle to assess results and guide further progress, particularly on non-verifiable taskshomogeneous recursive self-improvement (RSI), in which the model being improved (optimizee) also serves as its own optimizer and evaluator. The central question is how to obtain useful self-supervision within this homogeneous loop.

Human societies offer a useful analogy for addressing this question. Across foraging communities, farming settlements, industrial societies, and global information networks, people have expanded their capabilities in part by dividing work, combining ideas, and passing on what they learnThe Structure of Scientific Revolutions, Thomas Kuhn emphasized the shared examples through which scientists learn to recognize and solve problems

Inspired by this pattern, we ask whether a model can learn from its own team's collective intelligence, which an agent working alone may not reliably exhibit.

We introduce Multi-Agent Self-Supervision (MASS) to turn this idea into a learning loop. MASS uses the same model as the executor, optimizer, and evaluator. The optimizer first searches for the best team workflow using feedback from the evaluator, then trains the executor on executions from the self-judged best team. The updated model then fills all three roles, and the process repeats. Therefore, MASS provides a self-supervision signal for RSI by distilling the team's collective experience into the shared model.

We study homogeneous RSI, in which one shared language model serves as the task-solving agent, workflow optimizer, and evaluator:

Agent

The task solver, or optimizee. Produces a workspace containing results, plots, and a research report.

Workflow optimizer

Uses the feedback from the evaluator to revise how agents divide the task, exchange results, and check one another's work.

Evaluator

Compares two workspaces, chooses which better satisfies the task, and explains the judgment.

We use non-verifiable research tasks as training data. We use non-verifiable to mean that an automatic rule-based signal is not available and that only a self-evaluation signal (whose correctness is unknown) guides the overall learning loop.

To measure progress, we separately ask GPT-5.5 and Claude Opus 4.8 to assess the completed work. Their judgments are withheld from MASS. This gives us an outside measure of improvement in a process whose supervision comes from the model itself.

MASS uses workflow optimization to elicit collective intelligence, then distills the team's experience into the shared model. The inner loop optimizes the team's workflow while the model weights remain fixed. The model evaluates and judges for itself which workflow is optimal. The outer loop then updates those weights by learning from executions of the optimized workflow. The interactive diagram below follows both loops.

A workflow is a written plan for how a multi-agent team should work together. The orchestrator assigns work to subagents and integrates their results. The plan is added to the task prompt, making the team's organization available for the model to inspect and revise

For example, an evolved finance workflow arranges data preparation, feature construction, modeling, and backtesting into a sequence, followed by an audit. Its data contract calls for explicit checks on the timing of the data:

Evolved finance workflow, condensed from the paper

Call order
  Data → Features → Modeling → Backtest
       → Audit → Reconciliation → Report

Data specialist's output contract
  validation_report:
    no_future_leakage:           bool
    train_end_before_val_start:  bool
    val_end_before_test_start:   bool

A backtest can look convincing if it lets historical decisions use information from the future. This contract requires the data specialist to check for that error and report whether the time splits are valid. It also tells later agents what evidence they should receive before proceeding. Part of the research plan is thus expressed in what agents must establish and pass on to one another

For each training task, MASS retains a champion: the best workflow found so far according to the current model's judgments. Each iteration has three steps:

The evaluator judges the completed work; the optimizer uses that judgment to revise the process that produced it. Optimizer's system prompt emphasizes contracts and hop, directing workflow search toward what information agents exchange and when they receive it.

Even with an optimized workflow, the quality of long-horizon research work can vary from run to run. MASS therefore executes self-judged optimal workflows (or champion workflows, meaning the optimal workflows found so far) with multiple random seeds and uses the current model to rank the resulting workspaces, retaining the highest-ranked trajectories. Each trajectory records the messages, tool calls, and tool results of the orchestrator and its subagents.

Removing the workflow from the orchestrator's training prompt encourages the model to learn how to organize the work from the task itself.

We run two complete rounds of MASS starting from Qwen3.6-27B in the qwen-code coding-agent environment. Our synthetic research suite contains 12 tasks in finance, robotics, and pharmacy, each requiring code, quantitative results, and a research report

We compare three generations: the base model $\mathcal{L}^{(0)}$, the first-round model $\mathcal{L}^{(1)}$, and the second-round model $\mathcal{L}^{(2)}$. We examine the quality of their work, their ability to organize and judge it, and their use of coordination without a supplied workflow.

On the three synthetic test tasks, the win rate against the base rises from 53.9% after one round to 69.9% after two rounds (Table 1). The second-round model also wins 60.8% of comparisons against the first-round model.

Table 1. Synthetic task performance. External-judge win rates across eight training tasks and three tasks held out from post-training. All models receive task prompts without an optimized workflow.
Evaluation tasks Round 1 vs. base Round 2 vs. base Round 2 vs. round 1
--- --- --- ---
Training tasks 65.6% 82.5% 68.1%
Test tasks 53.9% 69.9% 60.8%

To test transfer beyond the synthetic research suite, we evaluate public benchmarks. After two rounds, score per output token reaches 1.2–1.6× the base model's level across MLR-Bench, DSBench, ScienceAgentBench, and AstaBench (Figure 3)

The gains are strongest on four research benchmarks resembling the training tasks. Score per output token is 0.96× the base value on Terminal-Bench 2.0 and 0.94× on SWE-bench Verified

We observe that MASS enables role generalization in RSI. We train the model only on task-solving trajectories and observe that the capabilities of the workflow optimizer and workspace evaluator also improve. This enables RSI loop to be sustainable.

Figure 4 shows optimizer can find successful workflows sooner as cycle goes by.

In fact, Figure 4 is technically not a very fair comparison, since $ L^{(0)}$ denotes a setting in which the base model (optimizer) updates the base model (optimizee), and $ L^{(1)}$ denotes a setting in which the first-round model (optimizer) updates the first-round model (optimizee). For a fair comparison, we give the first-round model's workflows to the base model (Table 2). We observe that the base model with first-round search workflows wins 64% of comparisons against the base model with base-search workflow.

Table 2. Workflow transfer. In the highlighted comparison, the same base model executes both sets of workflows. External judges assess the resulting workspaces.
Executor Workflow source Compared with Win rate
--- --- --- ---
Round 1 Round-1 search Base executing base-search workflows 80%
Base Round-1 search Base executing base-search workflows 64%
Base Round-1 search Base without a workflow 63%

The evalutor also move closer to those of the strong external model judges (more details in paper). Agreement with those external judges rises from 73% to 93% across the three cycles. Overall, these results suggest that task-solving examples can also improve capabilities used to produce and select the next round's training data.

What really makes an optimal multi-agent workflow effective?

We analyze the roles, instructions, contracts, and hops of the champion workflows (the best workflows found so far). For each component, we estimate how much its text reveals about the task. In the finance example above, a role such as “careful data scientist” could apply to many tasks. A contract specifying leakage checks for a trading backtest is much more distinctive

Figure 5 shows that between the first and final iterations of workflow optimization, an optimal multi-agent workflow tends to encode task information more strongly in contracts and hops and less in roles and instructions. We do not show this here, but mutual dependencies among the components generally decrease too (interesting!). This pattern suggests that the LLM optimizer evolves the workflow components to be mutually independent but optimizes the information flow via contracts and hops (see this blog post for deep, intuitive explanations). I would say that workflow optimization over discrete components resembles discovering a basis vector for that harness space.

We observe that trajectories from multi-agent systems can be more effective than trajectories from single agents in RSI. Multi-agent execution also changes the structure of the training data. The orchestrator's conversation shows how to divide a broad task into focused assignments and coordinate their execution. A subagent's conversation shows how to carry out an individual assignment. Specifically, we observe that jointly learning coordination from the orchestrator and bounded execution from subagents provides more effective supervision per training token than learning from long single-agent trajectories (Figure 6)

We compare students trained independently from the same base model on single-agent trajectories (S, S<sup>+</sup>, S<sup>++</sup>), multi-agent trajectories (M, M<sup>+</sup>), and a mixture (X). A plus sign denotes a larger training set; M<sup>+</sup> is the first-round MASS model. All points below report results pooled across four training tasks and one test task.

Against the base model, M<sup>+</sup> reaches a 68.3% win rate after processing 28.8M supervised training tokens, while S<sup>++</sup> reaches 64.0% after 42M tokens. The MASS model thus achieves the higher score against this reference with 31% fewer supervised tokens

Higher-quality trajectories may contribute to this advantage. In a separate analysis, strong external judges prefer the multi-agent systems' work in 85.1% of comparisons with single-agent outputs. Team runs also generate about 2.4× as many tokens on average

These results suggest that trajectories from multi-agent systems can be more effective than trajectories from single agents in RSI, extending their value beyond test-time scaling.

We observed that the bottleneck in RSI can be the optimizer, not the evaluator. We kept the base model as the task solver and swapped in a stronger external model (GPT-5.5) for the evaluator only, or for both the evaluator and the optimizer.

Figure 7 shows that upgrading only the evaluator helps a little, but upgrading the optimizer as well helps a lot. The strictly homogeneous self-improvement loop does eventually find winning workflows for 11 of 12 tasks; it just takes longer. Our interpretation is that the strong evaluator may provide high-quality feedback, but the weak optimizer is not capable enough to digest it. This suggests that, under a limited budget, one may have to prioritize the optimizer over the evaluator within an RSI loop unit.

Figure 8 also shows that workflow search is noisy. In text-space optimization, reaching a good workflow and improving on it stably are different problems. This suggests that the LLM optimizer may need an equivalent of a learning-rate scheduler to stabilize the search in text space.

In Tennyson's Ulysses, Odysseus imagines another voyage in pursuit of knowledge “Beyond the utmost bound of human thought.”Odyssean moment when continued learning must proceed without a stronger supervisor. The challenge concerns both who can supervise further learning and how progress should be judged. Human experts may be able to state a goal without specifying an objective function that fully captures it.

Writing is a good example: expert writers can recognize strong writing, yet struggle to express their judgment in a rubric that applies reliably across contexts. For RSI, the research question is how AI can turn incomplete human guidance into useful criteria for learning, and how to test whether those criteria continue to capture the intended goal as its capabilities grow. This becomes especially challenging when the model being improved also serves as its own optimizer and evaluator: the judgments that guide improvement must themselves remain open to correction.

Keeping a system's self-evaluation open to correction is also essential for safety. In a homogeneous loop, a shared blind spot could allow a flawed result to be produced, accepted as an improvement, and incorporated into the next model. If humans also struggle to recognize the flaw, successive updates may reinforce it. The safety concern is that the same weakness can compromise both the system's behavior and the mechanism meant to correct it.

In this sense, shared blind spots may be an Achilles' heel of recursive self-improvement. Understanding them matters for controlling a more capable system because it reveals where reliance on the system's own judgments is least justified. That knowledge could guide independent checks, constraints on autonomous actions, and conditions for intervention. Identifying a weakness alone does not establish control; we must also show that we can detect its consequences and intervene effectively as the system changes.

OpenAI's research on chain-of-thought monitorability illustrates one part of this challenge: reasoning traces can help reveal misbehavior, but they provide incomplete evidence

@misc{lee2026mass,
  title         = {Recursive Self-Improvement through Multi-Agent Self-Supervision},
  author        = {Lee, Hyunin and Xu, Jinglue and Seely, Jeffrey and Lee, Donghyun and Sojoudi, Somayeh and Zaharia, Matei and Tang, Yujin},
  year          = {2026},
  eprint        = {2610.12176},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  url           = {https://arxiv.org/abs/2610.12176}
}
── more in #artificial-intelligence 4 stories · sorted by recency
── more on @sakana ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/recursive-self-impro…] indexed:0 read:12min 2026-10-11 · —