cd/entity/SWE-bench· home› entities› SWE-bench
grep -l @swe-bench /news/*.json | wc -l → 106

SWE-bench

mentions 106 type Organization page 6/6 feed RSS

// recent coverage 106 mentions

19:57
2026-05-27
deepswe.datacurve.ai
ai-agents

DeepSWE Measuring frontier coding agents

DataCurve released DeepSWE, a new benchmark for evaluating frontier coding agents on original, long-horizon software engineering tasks. The benchmark features contamination-free tasks written from scr…

01:01
2026-05-20
dev.to
artificial-intelligence

DeepSeek V4 vs Claude Opus 4.5 for coding: benchmark comparison

Claude Opus 4.5 achieves an 80.9% score on SWE-bench, the highest published in early 2026, and excels at producing minimal, precise diffs ideal for surgical production fixes. DeepSeek V4 is stronger f…

← prev page 6 / 6
// co-occurs with top 8 entities
// topics top 6 topics