cd/entity/SWE-Bench Verified· home entities SWE-Bench Verified
grep -l @swe-bench verified /news/*.json | wc -l → 16

SWE-Bench Verified

mentions 16 type Person feed RSS

// recent coverage 16 mentions

18:40
2026-08-11
runtimewire.com
artificial-intelligence

Microsoft upgrades its Copilot coding model 10 weeks after launch

Microsoft released MAI-Code-1.1-Flash on August 11th and put the coding model into production in GitHub Copilot, replacing the first version of its in-house coding model just 10 weeks after launch. Th…

04:00
2026-07-21
machinebrief.com
artificial-intelligence

SWE-Pruner Pro: The Coder LLM Already Knows What to Prune

Researchers propose SWE-Pruner Pro, a method that prunes tool outputs directly inside a coding agent by using the agent's own internal representations to label each line as keep or prune, saving up to…

00:00
2026-07-19
rizz.dev
artificial-intelligence

Ways to Make the Most Out of Claude Fable 5 (2026)

Anthropic's Claude Fable 5, released June 9, 2026, leads the SWE-Bench Verified leaderboard with a 95% issue resolution rate, far ahead of Claude Opus 4.8 at 88.6%. The model's 1M-token context window…

23:09
2026-07-14
arxiv.org
large-language-models

LLM-as-a-Verifier: A General-Purpose Verification Framework

Researchers introduced LLM-as-a-Verifier, a general-purpose verification framework that computes continuous scores from scoring token logits to determine solution correctness without additional traini…

17:08
2026-07-04
byteiota.com
large-language-models

Claude Sonnet 5: What Developers Need to Know Before Migrating

Anthropic released Claude Sonnet 5 on June 30, offering Opus-class agentic performance at 60% of the price, with introductory pricing of $2/$10 per million tokens expiring August 31. The model beats O…

01:11
2026-06-30
byteiota.com
artificial-intelligence

Ornith 1.0 Beats Claude at Coding — Runs on One GPU

DeepReinforce AI released Ornith 1.0, an open-source coding model family that scores 82.4 on SWE-Bench Verified, outperforming Claude Opus 4.7's 80.8, and runs the 35B variant locally on a single RTX …

00:00
2026-06-11
telnyx.com
large-language-models

Kimi K2.6 Now Available for Telnyx AI Assistants

Telnyx has made Moonshot AI's Kimi K2.6 model available for AI Assistants in the US region, offering developers on-network inference to reduce latency and simplify infrastructure. The model scores 58.…

// co-occurs with top 8 entities
// topics top 6 topics