cd/entity/SWE-bench· home› entities› SWE-bench
grep -l @swe-bench /news/*.json | wc -l → 105

SWE-bench

mentions 105 type Organization page 2/6 feed RSS

// recent coverage 105 mentions

20:51
2026-08-24
tokenstead.ai
artificial-intelligence

Ornith-1.0-35B

Ornith AI released Ornith-1.0-35B, a 35B-parameter mixture-of-experts model with 3B active parameters per token, on June 25, 2026, under an MIT license on HuggingFace, featuring a 262K context window.…

18:37
2026-08-20
twitter.com
artificial-intelligence

When Building Gets Cheap, Knowing What to Build Gets Expensive

AI coding agents have made building software cheap, but AdaL's developers found that cheap implementation increases the risk of building the wrong thing. After building a benchmark first, they discove…

12:00
2026-08-20
kdnuggets.com
artificial-intelligence

Top 10 Open-Source Benchmarks for AI Coding Agents in 2026

SWE-bench remains the most widely used open-source benchmark for AI coding agents, with 2,294 tasks from 12 Python repositories, but newer benchmarks like Terminal-Bench, SWE-Bench Pro, and Senior SWE…

10:43
2026-08-16
aicharts.grok.me
artificial-intelligence

An AI agent can burn 100× the tokens of a chat turn

An AI agent session can consume 50,000–200,000 tokens, compared to a few hundred to a couple thousand for a chat turn, with multi-agent jobs exceeding one million tokens, according to illustrative ran…

12:11
2026-08-05
byteiota.com
artificial-intelligence

TencentDB Agent Memory Hits #1 GitHub: Fix Agent Amnesia

Tencent's open-source AI agent memory system, TencentDB Agent Memory, reached #1 on GitHub Trending this week, adding nearly 2,000 stars in a single day to reach 14,600 total. The tool provides persis…

22:07
2026-07-31
dev.to
large-language-models

Claude Sonnet 5 vs Opus 5: A Real-World Comparison (2026)

Anthropic's Claude Sonnet 5 and Claude Opus 5, released in mid-2026, offer distinct strengths and pricing trade-offs. Sonnet 5 excels at high-volume coding and content tasks with a 72.7% SWE-bench Ver…

14:00
2026-07-31
dev.to
developer-tools

My context selector beat grep. An agent with grep beat it.

A developer's context-selection tool, cognitive-cache, initially appeared no better than grep in a small benchmark, but a larger SWE-bench evaluation showed it significantly outperformed grep and TF-I…

← prev page 2 / 6 next →
// co-occurs with top 8 entities
// topics top 6 topics