cd/entity/SWE-bench· home entities SWE-bench
grep -l @swe-bench /news/*.json | wc -l → 82

SWE-bench

mentions 82 type Organization page 5/5 feed RSS

// recent coverage 82 mentions

19:57
2026-05-27
deepswe.datacurve.ai
ai-agents

DeepSWE Measuring frontier coding agents

DataCurve released DeepSWE, a new benchmark for evaluating frontier coding agents on original, long-horizon software engineering tasks. The benchmark features contamination-free tasks written from scr…

01:01
2026-05-20
dev.to
artificial-intelligence

DeepSeek V4 vs Claude Opus 4.5 for coding: benchmark comparison

Claude Opus 4.5 achieves an 80.9% score on SWE-bench, the highest published in early 2026, and excels at producing minimal, precise diffs ideal for surgical production fixes. DeepSeek V4 is stronger f…

← prev page 5 / 5
// co-occurs with top 8 entities
// topics top 6 topics