cd/entity/SWE-bench Verified· home› entities› SWE-bench Verified
grep -l @swe-bench verified /news/*.json | wc -l → 99

SWE-bench Verified

mentions 99 type Person page 1/5 feed RSS

// recent coverage 99 mentions

18:40
2026-10-08
yuzhenmao.github.io
ai-agents

DeLM: Decentralized Multi-Agent Systems with Shared Context

Researchers propose Decentralized Language Models (DeLM), a coordination layer that lets multiple LLM agents work asynchronously through a shared context and a task queue with no main agent in the loo…

04:00
2026-10-01
arxiv.org
ai-agents

TomasuLLM: Out-of-Order Speculative Execution for LLM Agents

Researchers released TomasuLLM, a runtime that executes LLM agent tool calls out of trajectory order while preserving task-execution correctness, according to arXiv paper 2609.38201v1. TomasuLLM draft…

20:19
2026-09-30
superml.dev
ai-agents

Agent Eval Costs: Stop Runs You Can Already Predict

A paper posted September 2, 2026, "EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction" (Shi, Sun, Dong, Wan, Lo, Gu), proposes halting agent benchmark runs whose outcome is already predi…

15:46
2026-09-28
dev.to
ai-agents

What one agent run actually costs

A developer's analysis of roughly 4,300 real coding-agent sessions (Claude Code and Codex, 43 developers, ~350,000 LLM steps) found that the median LLM step re-sends about 119,000 cached prefix tokens…

00:00
2026-09-28
mindstudio.ai
large-language-models

OrcaSAQ2 27B: Run Qwen3.8-27B in 12GB With 3-Bit Quantization

OrcaRouter released OrcaSAQ2 27B, a 3-bit mixed-precision quantization of Qwen3.8-27B that shrinks the checkpoint from 54GB in BF16 to 12.3GB, a 77.2% reduction, under Apache-2.0. The model card repor…

20:52
2026-09-24
dev.to
ai-agents

Harness Engineering 101: How Coding Agents Actually Work

A developer's analysis of coding agent architecture shows that swapping only the harness around a fixed model — keeping weights, tasks and context window constant — raised SWE-bench Verified bug-fixin…

03:43
2026-09-20
dev.to
large-language-models

Claude Sonnet 4.5 vs GPT-5: Claude Wins Coding

A head-to-head comparison finds Anthropic's Claude Sonnet 4.5 outperforms OpenAI's GPT-5 on coding and agentic benchmarks, scoring 77.2% on SWE-bench Verified versus GPT-5's 74.9%, and 50.0% versus 43…

06:57
2026-09-18
runtimewire.com
ai-research

Epoch AI audits 15 benchmarks, finds nine flawed

Epoch AI launched Benchmark Reviews on September 17th and classified nine of its first 15 reviewed benchmark versions as Flawed, four as Verified, and two as Not Enough Info, flagging defects in tests…

00:00
2026-09-17
mutagent.io
ai-agents

Meta-Evaluation: Which Tool Builds the Better Agent

Mutagent Helix published a meta-evaluation comparing its agent-building tool against hand-driving Claude Code across twenty agent benchmarks, measuring each calibration pass by dollar cost and number …

page 1 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics