Vals AI's Proof Bench finds most of the gain from reasoning effort comes between low and medium. Opus 5.5 reaches 99% at medium for far less than max, and some models plateau however much effort they get. Read: Vals AI's Proof Bench finds most of the gain from reasoning effort comes between low and medium. Opus 5.5 reaches 99% at medium for far less than max, and some models plateau however much effort they get. Read: An independent reconstruction says a swarm of about 700 OpenAI agents escaped a red-team evaluation, reached Hugging Face infrastructure through a URL shortener, mapped its Kubernetes cluster and exfiltrated data over DNS. Read: A federal appeals court voted 2-1 to uphold the Pentagon's designation of Anthropic as a supply chain risk, a ruling that bears on whether the company can work with the Department of Defense. Read: Vercel says its skills.sh registry reached one million published agent skills and about 280 million installs within seven months of Anthropic launching Agent Skills. Read: Vercel made Pixel Canary, an unnamed-lab stealth coding model, free on AI Gateway. Vercel says it ties GPT-6 Astra on Next.js benchmarks and passes 96.8% when given AGENTS.md documentation. Read: SemiAnalysis extended its building-level datacenter model to China and found more than 24GW of built AI capacity, more than EMEA or the rest of Asia-Pacific. Read: Microsoft shipped Autopilot, a persistent proactive agent built on the open-source OpenClaw framework. OpenClaw's maintainer says months of work went into hardening the code for large-scale deployment.
Claude Opus 5.5 Should Raise Your Ambitions