cd/entity/SWE-bench· home› entities› SWE-bench
grep -l @swe-bench /news/*.json | wc -l → 106

SWE-bench

mentions 106 type Organization page 5/6 feed RSS

// recent coverage 106 mentions

23:20
2026-06-25
letsdatascience.com
ai-tools

GitHub evaluates Copilot agentic harness performance

GitHub announced two harness-level improvements to Copilot agentic sessions—prompt caching achieving 94% cache hit rates and deferred tool loading—plus a new Auto model selection feature using its HyD…

12:30
2026-06-24
andrewjesson.com
ai-agents

When Does Data Help Automated Context Engineering?

Claude Code can improve other AI agents without training data in four of seven tested applications, performing as well as with data. Data helps only where Claude Code's prior knowledge of the task run…

00:00
2026-06-22
epics.tech
ai-policy

The Outcry Grows, but the Capital Keeps Flowing

Institutional backlash against AI is mounting across public sector, media, and education, with New York City council members urging a classroom AI pause and a German newsroom caught using AI to write …

06:51
2026-06-19
unsiloed.ai
developer-tools

Claude Code and Codex as one pipeline

A technical guide argues that developers should run both Claude Code and OpenAI Codex as a single pipeline rather than choosing one, based on two months of testing on large codebases. Benchmarks show …

02:35
2026-06-19
deepswe.datacurve.ai
ai-research

DeepSWE v1.1

DeepSWE v1.1 updates the benchmark for long-horizon engineering tasks with isolated verification and structured test reports, making results more reproducible and harder to game. Pass rates remain clo…

17:37
2026-06-18
mroczek.dev
developer-tools

The Token Compression Illusion: Why I'm Skeptical of RTK

RTK, a tool that compresses terminal output for LLM agents, claims to cut token usage by 60-90% but faces skepticism due to misleading savings metrics, silent failure risks, lack of accuracy benchmark…

14:00
2026-06-18
github.com
developer-tools

clawmark: open-source CLAUDE.md A/B Testing CLI tool

Clawmark, an open-source Rust CLI tool, enables A/B testing of CLAUDE.md files by evaluating two variants against five SWE-bench Lite tasks using Claude and Docker. The tool generates a comparison rep…

12:24
2026-06-05
sherwood.news
artificial-intelligence

Anthropic ponders self-improving AI

Anthropic reported that its AI model Claude now writes 80% of the company's internal code, raising questions about the potential for recursive self-improvement where AI systems autonomously enhance th…

20:09
2026-06-04
mendral.com
ai-agents

How we know if our agent is right

Mendral, an AI DevOps agent developer, cannot provide a single accuracy metric for its CI failure diagnosis agent despite processing 36,564 investigations across 5.7 million CI jobs and 14.4 billion l…

00:00
2026-05-29
deepresearch.ninja
artificial-intelligence

Claude Opus 4.8: Anatomy of Incremental Frontier Leadership

Anthropic released Claude Opus 4.8 on May 28, 2026, the eleventh major version in its Claude lineage, achieving a 69.2% score on SWE-bench Pro and introducing Dynamic Workflows for orchestrating hundr…

← prev page 5 / 6 next →
// co-occurs with top 8 entities
// topics top 6 topics