cd /news/artificial-intelligence/autonomous-repair-for-multi-agent-sy… · home topics artificial-intelligence article
[ARTICLE · art-84266] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Autonomous Repair for Multi-Agent Systems via Monte-Carlo Tree Search

Researchers propose MARS, a Monte Carlo Tree Search-based framework for automatically repairing multi-agent systems, and introduce StateMAS, a benchmark with 1,310 replayable failure trajectories. On StateMAS, MARS outperforms state-of-the-art methods by 3.0% to 12.1% absolute improvement across settings while maintaining comparable token consumption.

read1 min views1 publishedAug 3, 2026

arXiv:2607.29055v1 Announce Type: new Abstract: Multi-agent systems (MAS) are increasingly deployed to solve complex tasks. In case of incorrect or unsatisfactory outputs, users have to manually locate agent mistakes by inspecting agent trajectories (i.e., {\em failure attribution}) and provide feedback to refine the outputs (i.e., {\em repair}). Despite some recent work in MAS failure attribution, automated mechanisms to recover from such mistakes remain largely unexplored. To bridge this gap, we propose MARS, a search-based framework that formulates MAS repair as a Monte Carlo Tree Search (MCTS) process and navigates the vast space of potential repairs via diagnosis-guided expansion with taxonomy-augmented evaluation. Unlike standard MCTS, which evaluates a complete simulation via full rollout, MARS evaluates the agent trajectory using partial rollout to reduce token consumption. Furthermore, we introduce StateMAS, a large-scale MAS repair benchmark with 1,310 replayable multi-agent failure trajectories spanning four types of agent architectures and four LLM backbones. Experiments on StateMAS demonstrate that MARS consistently outperforms state-of-the-art methods, achieving an absolute improvement from 3.0% to 12.1% across all settings, while maintaining a comparable token consumption cost. The ablation study further confirms that taxonomy-augmented evaluation and diagnosis-guided expansion are critical to achieving these performance gains.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @mars 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/autonomous-repair-fo…] indexed:0 read:1min 2026-08-03 ·