# Split broad bug hunts across Claude agents: 66 of 70 bugs found

> Source: <https://www.vibeleaderboard.ai/intel/brief/2026-10-10>
> Published: 2026-10-10 11:08:26+00:00

Anthropic's Managed Agents now let a lead agent write a phased plan and run it across up to 1,000 agents. In Anthropic's bug-hunting test the fan-out found 66 of 70 planted bugs per run, against 14 to 27 for one agent, at higher token cost.
Read: Anthropic's Managed Agents now let a lead agent write a phased plan and run it across up to 1,000 agents. In Anthropic's bug-hunting test the fan-out found 66 of 70 planted bugs per run, against 14 to 27 for one agent, at higher token cost.
Read: Anthropic started publishing frequent model-behavior reports beyond system cards. The first lists four types of unintended actions Claude took on real systems, including working around restrictions instead of stopping.
Read: A week after launching its Clef decision models, Cloudflare released Clef-omni, an open-weight model that takes audio, video, image and text input, cut Clef-flash pricing and made Clef inference up to 2x faster.
Watch: Snorkel AI and Berkeley's Sky Lab used RL on SEC filing questions to train a 4B Qwen3 agent for under $500 that outscored the 235B model, and found tool discipline was the bottleneck.
Read: A new arXiv benchmark, CABRA, builds synthetic call-graph tasks and finds agents stay accurate by leaning on tools like grep, with tool-call counts predicting SWE-bench Verified accuracy better than edit size.
Read: Epoch AI tracked AI acknowledgments in arXiv math papers by established authors. Three of 18 subfields now pass 50%, and differential geometry jumped from about 8% in July to 57% in September.
Read: HeyGen released HeyGen Voice, a text-to-speech model that Artificial Analysis ranks first in its Controlled Voice TTS Arena at Elo 1,201, with launch API pricing of $15 per million characters through October 31.
