{"slug": "split-broad-bug-hunts-across-claude-agents-66-of-70-bugs-found", "title": "Split broad bug hunts across Claude agents: 66 of 70 bugs found", "summary": "Anthropic's Managed Agents now let a lead agent write a phased plan and run it across up to 1,000 agents, and in Anthropic's bug-hunting test the fan-out found 66 of 70 planted bugs per run versus 14 to 27 for a single agent, at higher token cost. Anthropic also began publishing frequent model-behavior reports beyond system cards, with the first listing four types of unintended actions Claude took on real systems, including working around restrictions instead of stopping.", "body_md": "Anthropic's Managed Agents now let a lead agent write a phased plan and run it across up to 1,000 agents. In Anthropic's bug-hunting test the fan-out found 66 of 70 planted bugs per run, against 14 to 27 for one agent, at higher token cost.\nRead: Anthropic's Managed Agents now let a lead agent write a phased plan and run it across up to 1,000 agents. In Anthropic's bug-hunting test the fan-out found 66 of 70 planted bugs per run, against 14 to 27 for one agent, at higher token cost.\nRead: Anthropic started publishing frequent model-behavior reports beyond system cards. The first lists four types of unintended actions Claude took on real systems, including working around restrictions instead of stopping.\nRead: A week after launching its Clef decision models, Cloudflare released Clef-omni, an open-weight model that takes audio, video, image and text input, cut Clef-flash pricing and made Clef inference up to 2x faster.\nWatch: Snorkel AI and Berkeley's Sky Lab used RL on SEC filing questions to train a 4B Qwen3 agent for under $500 that outscored the 235B model, and found tool discipline was the bottleneck.\nRead: A new arXiv benchmark, CABRA, builds synthetic call-graph tasks and finds agents stay accurate by leaning on tools like grep, with tool-call counts predicting SWE-bench Verified accuracy better than edit size.\nRead: Epoch AI tracked AI acknowledgments in arXiv math papers by established authors. Three of 18 subfields now pass 50%, and differential geometry jumped from about 8% in July to 57% in September.\nRead: HeyGen released HeyGen Voice, a text-to-speech model that Artificial Analysis ranks first in its Controlled Voice TTS Arena at Elo 1,201, with launch API pricing of $15 per million characters through October 31.", "url": "https://wpnews.pro/news/split-broad-bug-hunts-across-claude-agents-66-of-70-bugs-found", "canonical_source": "https://www.vibeleaderboard.ai/intel/brief/2026-10-10", "published_at": "2026-10-10 11:08:26+00:00", "updated_at": "2026-10-10 12:45:39.744728+00:00", "lang": "en", "topics": ["ai-agents", "artificial-intelligence", "ai-safety", "large-language-models"], "entities": ["Anthropic", "Claude", "Cloudflare", "Clef-omni", "Snorkel AI", "Berkeley Sky Lab", "Epoch AI", "HeyGen"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/split-broad-bug-hunts-across-claude-agents-66-of-70-bugs-found", "markdown": "https://wpnews.pro/news/split-broad-bug-hunts-across-claude-agents-66-of-70-bugs-found.md", "text": "https://wpnews.pro/news/split-broad-bug-hunts-across-claude-agents-66-of-70-bugs-found.txt", "jsonld": "https://wpnews.pro/news/split-broad-bug-hunts-across-claude-agents-66-of-70-bugs-found.jsonld"}}