cd/entity/SWE-bench· home› entities› SWE-bench
grep -l @swe-bench /news/*.json | wc -l → 106

SWE-bench

mentions 106 type Organization page 3/6 feed RSS

// recent coverage 106 mentions

14:00
2026-07-31
dev.to
developer-tools

My context selector beat grep. An agent with grep beat it.

A developer's context-selection tool, cognitive-cache, initially appeared no better than grep in a small benchmark, but a larger SWE-bench evaluation showed it significantly outperformed grep and TF-I…

17:56
2026-07-29
dev.to
artificial-intelligence

Model + Harness = Agent: The Gap Isn’t Where You Think

A developer reports that the gap between AI models and their agentic performance is often determined by the harness—the system around the model—rather than the model itself. Evidence from Moonshot's K…

19:40
2026-07-26
promptcube3.com
developer-tools

Copilot vs Raw API: What are you actually paying for?

GitHub Copilot's value lies not in raw model access but in its integrated harness that connects editor, terminal, and PR flow, according to a technical analysis. GitHub's evaluations across SWE-bench …

01:05
2026-07-26
promptcube3.com
large-language-models

Kimi K3 vs GPT-5 vs Claude 4 Opus: 2026 Comparison

Kimi K3 is 30x cheaper than GPT-5 for output tokens while leading the LMArena leaderboard with 1,289 ELO, according to a 2026 comparison. For a customer support bot handling 10M input and 5M output to…

05:09
2026-07-23
dev.to
artificial-intelligence

AI Weekly: MCP Goes Stateless, AMD Ships 2nm Silicon

The AI industry is shifting focus from raw novelty to durable interfaces, with the Model Context Protocol locking its largest revision ahead of a July 28 release, AMD unveiling its first 2nm x86 serve…

17:07
2026-07-22
dev.to
ai-agents

Agentic Code Security: What Autonomous AI Gets Wrong

BrassCoders finds that autonomous coding agents, such as Claude Code and SWE-agents, bypass human code review, allowing security vulnerabilities like hardcoded credentials to go undetected. The compan…

22:35
2026-07-21
fireworks.ai
artificial-intelligence

Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

Kimi K3, an open model from Moonshot AI, achieves near-parity with the closed-source Fable 5 on a suite of ~1,030 agentic tasks while costing a fraction of the price, according to a benchmark study by…

18:17
2026-07-21
khola.blog
artificial-intelligence

The Top-Down Bet Needs A Bottom-Up Audit

A mid-2026 audit of top-down AI-assisted software engineering shows agents absorbing implementation work on schedule, with SWE-bench Verified scores rising from 1.96% in October 2023 to near saturatio…

16:44
2026-07-20
i-programmer.info
ai-agents

DeepSWE - Best Benchmark For Evaluating AI Coding Agents?

DeepSWE, a new benchmarking platform for AI coding agents, achieves clearer separation between frontier models than older benchmarks like SWE-bench by using contamination-free tasks across 91 reposito…

04:14
2026-07-20
byteiota.com
artificial-intelligence

mini-swe-agent: 100 Lines That Beat Claude Code on SWE-bench

Princeton and Stanford researchers released mini-swe-agent, a coding agent in roughly 100 lines of Python that scores over 74% on SWE-bench Verified, beating Claude Code's scaffolding by six percentag…

00:00
2026-07-19
zackproser.com
artificial-intelligence

The Benchmark

A benchmark score is a manufactured number that passes through a chain of choices—sampling, prompting, scoring, and aggregation—each of which can change the final result while model weights stay fixed…

13:09
2026-07-17
byteiota.com
artificial-intelligence

Grok 4.5: Cursor-Trained Model at $2 Per Million Tokens

XAI released Grok 4.5 on July 8, trained on trillions of Cursor session tokens capturing real developer-agent interactions, achieving #1 on TAU-bench agentic tool use (71%) and resolving SWE-bench tas…

01:54
2026-07-16
dev.to
large-language-models

LLM Evals For Developer Tools: Useful, Correct, Safe

A developer argues that LLM evaluations for developer tools must measure three distinct axes: usefulness, correctness, and safety. Unlike chatbot evals, dev-tool evals rely on ground-truth signals lik…

← prev page 3 / 6 next →
// co-occurs with top 8 entities
// topics top 6 topics