cd/entity/SWE-bench· home entities SWE-bench
grep -l @swe-bench /news/*.json | wc -l → 82

SWE-bench

mentions 82 type Organization page 4/5 feed RSS

// recent coverage 82 mentions

06:51
2026-06-19
unsiloed.ai
developer-tools

Claude Code and Codex as one pipeline

A technical guide argues that developers should run both Claude Code and OpenAI Codex as a single pipeline rather than choosing one, based on two months of testing on large codebases. Benchmarks show …

02:35
2026-06-19
deepswe.datacurve.ai
ai-research

DeepSWE v1.1

DeepSWE v1.1 updates the benchmark for long-horizon engineering tasks with isolated verification and structured test reports, making results more reproducible and harder to game. Pass rates remain clo…

17:37
2026-06-18
mroczek.dev
developer-tools

The Token Compression Illusion: Why I'm Skeptical of RTK

RTK, a tool that compresses terminal output for LLM agents, claims to cut token usage by 60-90% but faces skepticism due to misleading savings metrics, silent failure risks, lack of accuracy benchmark…

14:00
2026-06-18
github.com
developer-tools

clawmark: open-source CLAUDE.md A/B Testing CLI tool

Clawmark, an open-source Rust CLI tool, enables A/B testing of CLAUDE.md files by evaluating two variants against five SWE-bench Lite tasks using Claude and Docker. The tool generates a comparison rep…

12:24
2026-06-05
sherwood.news
artificial-intelligence

Anthropic ponders self-improving AI

Anthropic reported that its AI model Claude now writes 80% of the company's internal code, raising questions about the potential for recursive self-improvement where AI systems autonomously enhance th…

20:09
2026-06-04
mendral.com
ai-agents

How we know if our agent is right

Mendral, an AI DevOps agent developer, cannot provide a single accuracy metric for its CI failure diagnosis agent despite processing 36,564 investigations across 5.7 million CI jobs and 14.4 billion l…

00:00
2026-05-29
deepresearch.ninja
artificial-intelligence

Claude Opus 4.8: Anatomy of Incremental Frontier Leadership

Anthropic released Claude Opus 4.8 on May 28, 2026, the eleventh major version in its Claude lineage, achieving a 69.2% score on SWE-bench Pro and introducing Dynamic Workflows for orchestrating hundr…

← prev page 4 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics