Give agents the bug report: unprompted, they fix under 5% of bugs SWE-sweep tested autonomous bug fixing by asking models to find and fix real bugs across 100 repositories with no ticket to guide them, and the best configuration fixed under 5% of bugs. The result indicates autonomous bug hunting still requires a human to specify what is wrong. SWE-sweep asked models to find and fix real bugs across 100 repositories with no ticket to guide them. The best configuration fixed under 5%, so autonomous bug hunting still needs a human to say what is wrong. Read: SWE-sweep asked models to find and fix real bugs across 100 repositories with no ticket to guide them. The best configuration fixed under 5%, so autonomous bug hunting still needs a human to say what is wrong. Read: Anthropic's Claude Code team introduced mods, TypeScript extensions shipped inside plugins that can alter behavior, customize the interface and add features. Read: Earendil released Pi 1.0, a minimal agent harness that adds Codemode, deferred tool loading and Anthropic cache warming, alongside Pi Durable, an experimental base for long-running agents that checkpoints state. Read: Cloudflare released Clef and Clef-flash, open-weight models for classification and routing with bounded structured outputs, hosted on Workers AI, plus a platform to fine-tune them with reinforcement learning. Read: DeepSeek released packaged desktop builds of DeepSeek Harness v0.2 preview for macOS and Windows. Linux users install it through the dsh npm package. Read: A new arXiv preprint defines bounded agent loops as a worker, a gate it cannot write to, and a declared budget, and audits shipped loops for gates that pass without checking anything. Read: Ethan Mollick revised his earlier view that managing agents requires careful organizational design, arguing the Bitter Lesson now applies to agent scaffolding.