cd /news/ai-agents/root-cause-attribution-is-a-search-p… · home topics ai-agents article
[ARTICLE · art-129826] src=machinebrief.com ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures

A new arXiv paper introduces Continual Search, an iterative framework for automated root-cause attribution (RCA) in long-horizon AI agent failures, reporting that it improves GPT-5.5's F1 score by more than 40%, from 0.349 to 0.498, on the newly introduced MegaRCA-Mix benchmark. The authors argue existing one-shot LLM judges settle on a plausible diagnosis early and leave critical evidence in longer execution traces unexamined, reducing RCA to a massive search problem. MegaRCA-Mix provides 50 human-annotated failure trials spanning long-horizon, execution-heavy tasks, and the paper reports that within the same model family lower-tier models can surpass higher-tier counterparts, showing effective search supersedes raw model scale.

by read1 min views2 publishedSep 15, 2026

arXiv:2609.13463v1 Announce Type: new Abstract: The increasing deployment of AI agents in long-horizon tasks yields massive execution logs. Diagnosing failures within these records is crucial for reliability, as it transforms outcome-level signals into actionable interventions. The sheer scale of the data renders human review impractical, driving the need for automated root-cause attribution (RCA). However, automated RCA methods using LLMs suffer from low diagnostic accuracy, especially as execution traces grow larger. They struggle because relevant information is often sparse, distributed across distant actions, and disconnected from the visible failure, reducing root-cause attribution to a massive search problem. Existing RCA methods typically rely on one-shot LLM judgments to diagnose failures from execution traces. While effective for shorter trajectories, these judges tend to settle on a plausible diagnosis early, leaving critical evidence in longer traces unexamined. We introduce Continual Search, an iterative framework that nudges the judge, over successive turns, to keep searching for unresolved diagnostic evidence. We evaluate Continual Search across four existing RCA benchmarks. Recognizing the lack of massive execution traces in current benchmarks, we introduce MegaRCA-Mix to evaluate RCA at scale. MegaRCA-Mix provides a challenging testbed of 50 human-annotated failure trials spanning long-horizon, execution-heavy tasks. Across multiple benchmark suites and model families, Continual Search consistently improves attribution performance. On MegaRCA-Mix, for example, it improves GPT-5.5's F1 score by more than 40%, from $0.349$ to $0.498$. More interestingly, within the same model family, lower-tier models can even surpass their higher-tier counterparts, demonstrating that effective search supersedes raw model scale.

── more in #ai-agents 4 stories · sorted by recency
── more on @continual search 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/root-cause-attributi…] indexed:0 read:1min 2026-09-15 ·