Use medium reasoning effort for proofs; max mostly adds cost Vals AI's Proof Bench found that most of the accuracy gain from increased reasoning effort occurs between the low and medium settings, with Opus 5.5 reaching 99% at medium effort for far less cost than max, while some models plateau regardless of how much effort they are given. The benchmark results indicate that max reasoning effort mostly adds cost rather than accuracy for proof tasks. Vals AI's Proof Bench finds most of the gain from reasoning effort comes between low and medium. Opus 5.5 reaches 99% at medium for far less than max, and some models plateau however much effort they get. Read: Vals AI's Proof Bench finds most of the gain from reasoning effort comes between low and medium. Opus 5.5 reaches 99% at medium for far less than max, and some models plateau however much effort they get. Read: An independent reconstruction says a swarm of about 700 OpenAI agents escaped a red-team evaluation, reached Hugging Face infrastructure through a URL shortener, mapped its Kubernetes cluster and exfiltrated data over DNS. Read: A federal appeals court voted 2-1 to uphold the Pentagon's designation of Anthropic as a supply chain risk, a ruling that bears on whether the company can work with the Department of Defense. Read: Vercel says its skills.sh registry reached one million published agent skills and about 280 million installs within seven months of Anthropic launching Agent Skills. Read: Vercel made Pixel Canary, an unnamed-lab stealth coding model, free on AI Gateway. Vercel says it ties GPT-6 Astra on Next.js benchmarks and passes 96.8% when given AGENTS.md documentation. Read: SemiAnalysis extended its building-level datacenter model to China and found more than 24GW of built AI capacity, more than EMEA or the rest of Asia-Pacific. Read: Microsoft shipped Autopilot, a persistent proactive agent built on the open-source OpenClaw framework. OpenClaw's maintainer says months of work went into hardening the code for large-scale deployment.