cd/entity/SWE-bench Verified· home› entities› SWE-bench Verified
grep -l @swe-bench verified /news/*.json | wc -l → 99

SWE-bench Verified

mentions 99 type Person page 2/5 feed RSS

// recent coverage 99 mentions

00:00
2026-09-11
digitalapplied.com
ai-safety

AI Benchmark Contamination: What Evidence Is Enough?

Digital Applied published an evidence-classification method on September 12, 2026, for assessing AI benchmark contamination claims, distinguishing training exposure, runtime answer access, and defecti…

10:14
2026-09-09
byteiota.com
ai-agents

OpenHands 1.0: Self-Hosted Coding Agent With Safety Sandbox

OpenHands 1.0, an open-source autonomous coding agent from All Hands AI, shipped this week with a production-grade Docker security sandbox, a built-in LLM-based security analyzer that rates actions LO…

03:01
2026-09-07
dev.to
ai-agents

Why Your AI Agent Should Just Be a Simple while Loop

A developer argues that most production AI agents should be built as simple while loops rather than complex frameworks like LangChain or CrewAI. The post highlights that top-performing agents on SWE-b…

07:22
2026-09-03
dev.to
ai-agents

False Completion Is the Real Failure Mode of Coding Agents

A developer argues that false completion, where AI coding agents report success while delivering incomplete or incorrect products, is the real failure mode in autonomous software development. The deve…

22:00
2026-08-31
ianreppel.org
artificial-intelligence

The Last Vendor

OpenAI's GPT-5.2, released in August 2026, achieves 95.4% on SWE-bench Verified, surpassing Claude Opus 5 by 0.6 points and becoming the top open-weights model, according to the Vals AI leaderboard. T…

19:37
2026-08-31
arxiv.org
artificial-intelligence

AI Agents Are Fundamentally Restructuring the Software Paradigm

A paper submitted to arXiv on June 4, 2026, and revised June 10, 2026, argues that AI agents—systems where large language models serve as the primary reasoning engine—constitute a fundamental restruct…

19:54
2026-08-28
techstrong.ai
large-language-models

IBM Granite 4.2 Adds a Reasoning Mode and Deeper Agent Training

IBM released Granite 4.2, an update to its open-weight 3B, 8B, and 30B language models, adding an optional reasoning mode and deeper agentic training for the 8B and 30B models. The models, available u…

02:07
2026-08-28
arxiv.org
ai-safety

Evomal: Self-Poisoning in Self-Evolving Coding Agents

A new arXiv paper (submitted Aug 26, 2026) reveals that self-evolving LLM coding agents can be poisoned through a self-propagating worm attack called EvoMal, which exploits the agents' habit of imitat…

10:33
2026-08-27
dev.to
artificial-intelligence

Your test agent isn't bad at clicking. It's bad at judging.

A developer from DevAssure argues that browser-based AI agents are failing at judgment, not actuation, citing benchmarks showing judges disagree with humans a third of the time and flawed ground truth…

22:25
2026-08-26
news.ycombinator.com
artificial-intelligence

I estimate reading code costs 2.1x more than writing it

A developer estimates that human code review costs $0.243 per changed line, 2.1 times the $0.114 model spend for an AI agent to produce an accepted changed line, based on a SmartBear/Cisco study, U.S.…

18:44
2026-08-26
arxiv.org
ai-agents

Agent Harness Evolution Shapes Coding Agent Quality

A controlled longitudinal study of 35 sequential releases of the Qwen Code CLI, holding the underlying LLM constant, found that agent harness evolution significantly impacts coding agent quality, with…

← prev page 2 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics