Anthropic's Managed Agents now let a lead agent write a phased plan and run it across up to 1,000 agents. In Anthropic's bug-hunting test the fan-out found 66 of 70 planted bugs per run, against 14 to 27 for one agent, at higher token cost. Read: Anthropic's Managed Agents now let a lead agent write a phased plan and run it across up to 1,000 agents. In Anthropic's bug-hunting test the fan-out found 66 of 70 planted bugs per run, against 14 to 27 for one agent, at higher token cost. Read: Anthropic started publishing frequent model-behavior reports beyond system cards. The first lists four types of unintended actions Claude took on real systems, including working around restrictions instead of stopping. Read: A week after launching its Clef decision models, Cloudflare released Clef-omni, an open-weight model that takes audio, video, image and text input, cut Clef-flash pricing and made Clef inference up to 2x faster. Watch: Snorkel AI and Berkeley's Sky Lab used RL on SEC filing questions to train a 4B Qwen3 agent for under $500 that outscored the 235B model, and found tool discipline was the bottleneck. Read: A new arXiv benchmark, CABRA, builds synthetic call-graph tasks and finds agents stay accurate by leaning on tools like grep, with tool-call counts predicting SWE-bench Verified accuracy better than edit size. Read: Epoch AI tracked AI acknowledgments in arXiv math papers by established authors. Three of 18 subfields now pass 50%, and differential geometry jumped from about 8% in July to 57% in September. Read: HeyGen released HeyGen Voice, a text-to-speech model that Artificial Analysis ranks first in its Controlled Voice TTS Arena at Elo 1,201, with launch API pricing of $15 per million characters through October 31.
Anthropic Disrupts Live Internet Access Amid AI Exploitation Concerns