cd/entity/SWE-bench· home entities SWE-bench
grep -l @swe-bench /news/*.json | wc -l → 82

SWE-bench

mentions 82 type Organization page 2/5 feed RSS

// recent coverage 82 mentions

19:40
2026-07-26
promptcube3.com
developer-tools

Copilot vs Raw API: What are you actually paying for?

GitHub Copilot's value lies not in raw model access but in its integrated harness that connects editor, terminal, and PR flow, according to a technical analysis. GitHub's evaluations across SWE-bench …

01:05
2026-07-26
promptcube3.com
large-language-models

Kimi K3 vs GPT-5 vs Claude 4 Opus: 2026 Comparison

Kimi K3 is 30x cheaper than GPT-5 for output tokens while leading the LMArena leaderboard with 1,289 ELO, according to a 2026 comparison. For a customer support bot handling 10M input and 5M output to…

05:09
2026-07-23
dev.to
artificial-intelligence

AI Weekly: MCP Goes Stateless, AMD Ships 2nm Silicon

The AI industry is shifting focus from raw novelty to durable interfaces, with the Model Context Protocol locking its largest revision ahead of a July 28 release, AMD unveiling its first 2nm x86 serve…

17:07
2026-07-22
dev.to
ai-agents

Agentic Code Security: What Autonomous AI Gets Wrong

BrassCoders finds that autonomous coding agents, such as Claude Code and SWE-agents, bypass human code review, allowing security vulnerabilities like hardcoded credentials to go undetected. The compan…

22:35
2026-07-21
fireworks.ai
artificial-intelligence

Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

Kimi K3, an open model from Moonshot AI, achieves near-parity with the closed-source Fable 5 on a suite of ~1,030 agentic tasks while costing a fraction of the price, according to a benchmark study by…

18:17
2026-07-21
khola.blog
artificial-intelligence

The Top-Down Bet Needs A Bottom-Up Audit

A mid-2026 audit of top-down AI-assisted software engineering shows agents absorbing implementation work on schedule, with SWE-bench Verified scores rising from 1.96% in October 2023 to near saturatio…

16:44
2026-07-20
i-programmer.info
ai-agents

DeepSWE - Best Benchmark For Evaluating AI Coding Agents?

DeepSWE, a new benchmarking platform for AI coding agents, achieves clearer separation between frontier models than older benchmarks like SWE-bench by using contamination-free tasks across 91 reposito…

04:14
2026-07-20
byteiota.com
artificial-intelligence

mini-swe-agent: 100 Lines That Beat Claude Code on SWE-bench

Princeton and Stanford researchers released mini-swe-agent, a coding agent in roughly 100 lines of Python that scores over 74% on SWE-bench Verified, beating Claude Code's scaffolding by six percentag…

00:00
2026-07-19
zackproser.com
artificial-intelligence

The Benchmark

A benchmark score is a manufactured number that passes through a chain of choices—sampling, prompting, scoring, and aggregation—each of which can change the final result while model weights stay fixed…

13:09
2026-07-17
byteiota.com
artificial-intelligence

Grok 4.5: Cursor-Trained Model at $2 Per Million Tokens

XAI released Grok 4.5 on July 8, trained on trillions of Cursor session tokens capturing real developer-agent interactions, achieving #1 on TAU-bench agentic tool use (71%) and resolving SWE-bench tas…

01:54
2026-07-16
dev.to
large-language-models

LLM Evals For Developer Tools: Useful, Correct, Safe

A developer argues that LLM evaluations for developer tools must measure three distinct axes: usefulness, correctness, and safety. Unlike chatbot evals, dev-tool evals rely on ground-truth signals lik…

07:39
2026-07-15
machinebrief.com
artificial-intelligence

AI Benchmarks: When Is Enough Truly Enough?

A study examining partial evaluations on AI agent benchmarks including SWE-bench, AppWorld, and tau-bench finds that partial budgets are only valid when they replicate the full benchmark's final decis…

04:00
2026-07-15
arxiv.org
artificial-intelligence

Token Reduction Is Not Cost Reduction

A new study from arXiv:2607.12161v1 analyzing 2,848 Claude Code runs across 103 tasks finds that reducing retrieved context or tool output does not reliably lower billed costs for API-based coding age…

← prev page 2 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics