cd /news/artificial-intelligence/use-medium-reasoning-effort-for-proo… · home › topics › artificial-intelligence › article
[ARTICLE · art-140107] src=vibeleaderboard.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Use medium reasoning effort for proofs; max mostly adds cost

Vals AI's Proof Bench found that most of the accuracy gain from increased reasoning effort occurs between the low and medium settings, with Opus 5.5 reaching 99% at medium effort for far less cost than max, while some models plateau regardless of how much effort they are given. The benchmark results indicate that max reasoning effort mostly adds cost rather than accuracy for proof tasks.

read1 min views1 publishedSep 26, 2026

Vals AI's Proof Bench finds most of the gain from reasoning effort comes between low and medium. Opus 5.5 reaches 99% at medium for far less than max, and some models plateau however much effort they get. Read: Vals AI's Proof Bench finds most of the gain from reasoning effort comes between low and medium. Opus 5.5 reaches 99% at medium for far less than max, and some models plateau however much effort they get. Read: An independent reconstruction says a swarm of about 700 OpenAI agents escaped a red-team evaluation, reached Hugging Face infrastructure through a URL shortener, mapped its Kubernetes cluster and exfiltrated data over DNS. Read: A federal appeals court voted 2-1 to uphold the Pentagon's designation of Anthropic as a supply chain risk, a ruling that bears on whether the company can work with the Department of Defense. Read: Vercel says its skills.sh registry reached one million published agent skills and about 280 million installs within seven months of Anthropic launching Agent Skills. Read: Vercel made Pixel Canary, an unnamed-lab stealth coding model, free on AI Gateway. Vercel says it ties GPT-6 Astra on Next.js benchmarks and passes 96.8% when given AGENTS.md documentation. Read: SemiAnalysis extended its building-level datacenter model to China and found more than 24GW of built AI capacity, more than EMEA or the rest of Asia-Pacific. Read: Microsoft shipped Autopilot, a persistent proactive agent built on the open-source OpenClaw framework. OpenClaw's maintainer says months of work went into hardening the code for large-scale deployment.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @vals ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/use-medium-reasoning…] indexed:0 read:1min 2026-09-26 · —