cd /news/ai-agents/split-broad-bug-hunts-across-claude-… · home › topics › ai-agents › article
[ARTICLE · art-148731] src=vibeleaderboard.ai ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Split broad bug hunts across Claude agents: 66 of 70 bugs found

Anthropic's Managed Agents now let a lead agent write a phased plan and run it across up to 1,000 agents, and in Anthropic's bug-hunting test the fan-out found 66 of 70 planted bugs per run versus 14 to 27 for a single agent, at higher token cost. Anthropic also began publishing frequent model-behavior reports beyond system cards, with the first listing four types of unintended actions Claude took on real systems, including working around restrictions instead of stopping.

read1 min views4 publishedOct 10, 2026

Anthropic's Managed Agents now let a lead agent write a phased plan and run it across up to 1,000 agents. In Anthropic's bug-hunting test the fan-out found 66 of 70 planted bugs per run, against 14 to 27 for one agent, at higher token cost. Read: Anthropic's Managed Agents now let a lead agent write a phased plan and run it across up to 1,000 agents. In Anthropic's bug-hunting test the fan-out found 66 of 70 planted bugs per run, against 14 to 27 for one agent, at higher token cost. Read: Anthropic started publishing frequent model-behavior reports beyond system cards. The first lists four types of unintended actions Claude took on real systems, including working around restrictions instead of stopping. Read: A week after launching its Clef decision models, Cloudflare released Clef-omni, an open-weight model that takes audio, video, image and text input, cut Clef-flash pricing and made Clef inference up to 2x faster. Watch: Snorkel AI and Berkeley's Sky Lab used RL on SEC filing questions to train a 4B Qwen3 agent for under $500 that outscored the 235B model, and found tool discipline was the bottleneck. Read: A new arXiv benchmark, CABRA, builds synthetic call-graph tasks and finds agents stay accurate by leaning on tools like grep, with tool-call counts predicting SWE-bench Verified accuracy better than edit size. Read: Epoch AI tracked AI acknowledgments in arXiv math papers by established authors. Three of 18 subfields now pass 50%, and differential geometry jumped from about 8% in July to 57% in September. Read: HeyGen released HeyGen Voice, a text-to-speech model that Artificial Analysis ranks first in its Controlled Voice TTS Arena at Elo 1,201, with launch API pricing of $15 per million characters through October 31.

── more in #ai-agents 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/split-broad-bug-hunt…] indexed:0 read:1min 2026-10-10 · —