cd/entity/SWE-rebenchยท homeโ€บ entitiesโ€บ SWE-rebench
grep -l @swe-rebench /news/*.json | wc -l โ†’ 2

SWE-rebench

mentions 2 type Organization feed RSS

// recent coverage 2 mentions

17:09
2026-07-31
sourcefeed.dev
artificial-intelligence

Frontier Coding Models Are Two Points Apart. Costs Aren't

SWE-rebench's latest leaderboard shows frontier coding models within 2.2 points of each other on resolved rate, with Fable 5 (Anthropic) at 64.5%, Grok 4.5 (xAI) at 63.8%, Opus 5 (Anthropic) at 63.4%,โ€ฆ

22:32
2026-06-17
transluce.org
ai-agents

A framework for verifiable analysis of AI behavior

Transluce Labs released Docent, a framework for verifiable analysis of AI behavior that allows humans to inspect and audit every step of an AI agent's analysis. The framework uses analysis plans with โ€ฆ

// co-occurs with top 8 entities
// topics top 6 topics