cd/entity/SWE-bench· home entities SWE-bench
grep -l @swe-bench /news/*.json | wc -l → 82

SWE-bench

mentions 82 type Organization page 1/5 feed RSS

// recent coverage 82 mentions

18:37
2026-08-20
twitter.com
artificial-intelligence

When Building Gets Cheap, Knowing What to Build Gets Expensive

AI coding agents have made building software cheap, but AdaL's developers found that cheap implementation increases the risk of building the wrong thing. After building a benchmark first, they discove…

12:00
2026-08-20
kdnuggets.com
artificial-intelligence

Top 10 Open-Source Benchmarks for AI Coding Agents in 2026

SWE-bench remains the most widely used open-source benchmark for AI coding agents, with 2,294 tasks from 12 Python repositories, but newer benchmarks like Terminal-Bench, SWE-Bench Pro, and Senior SWE…

10:43
2026-08-16
aicharts.grok.me
artificial-intelligence

An AI agent can burn 100× the tokens of a chat turn

An AI agent session can consume 50,000–200,000 tokens, compared to a few hundred to a couple thousand for a chat turn, with multi-agent jobs exceeding one million tokens, according to illustrative ran…

12:11
2026-08-05
byteiota.com
artificial-intelligence

TencentDB Agent Memory Hits #1 GitHub: Fix Agent Amnesia

Tencent's open-source AI agent memory system, TencentDB Agent Memory, reached #1 on GitHub Trending this week, adding nearly 2,000 stars in a single day to reach 14,600 total. The tool provides persis…

22:07
2026-07-31
dev.to
large-language-models

Claude Sonnet 5 vs Opus 5: A Real-World Comparison (2026)

Anthropic's Claude Sonnet 5 and Claude Opus 5, released in mid-2026, offer distinct strengths and pricing trade-offs. Sonnet 5 excels at high-volume coding and content tasks with a 72.7% SWE-bench Ver…

14:00
2026-07-31
dev.to
developer-tools

My context selector beat grep. An agent with grep beat it.

A developer's context-selection tool, cognitive-cache, initially appeared no better than grep in a small benchmark, but a larger SWE-bench evaluation showed it significantly outperformed grep and TF-I…

17:56
2026-07-29
dev.to
artificial-intelligence

Model + Harness = Agent: The Gap Isn’t Where You Think

A developer reports that the gap between AI models and their agentic performance is often determined by the harness—the system around the model—rather than the model itself. Evidence from Moonshot's K…

page 1 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics