cd/entity/SWE-Bench Pro· home entities SWE-Bench Pro
grep -l @swe-bench pro /news/*.json | wc -l → 39

SWE-Bench Pro

mentions 39 type Person page 1/2 feed RSS

// recent coverage 39 mentions

04:14
2026-08-27
byteiota.com
developer-tools

VS Code 1.135: Agent Host Protocol Ships, Sessions Go Portable

VS Code 1.135, released August 26, introduces the Agent Host Protocol (AHP), an open, MIT-licensed specification that decouples AI agent sessions from the editor window, allowing sessions to persist i…

12:00
2026-08-20
kdnuggets.com
artificial-intelligence

Top 10 Open-Source Benchmarks for AI Coding Agents in 2026

SWE-bench remains the most widely used open-source benchmark for AI coding agents, with 2,294 tasks from 12 Python repositories, but newer benchmarks like Terminal-Bench, SWE-Bench Pro, and Senior SWE…

09:01
2026-08-19
glad-ia-tor.com
artificial-intelligence

Claude Sonnet 5 vs Opus 4.8: When the $2 Model Beats the $25 One

Anthropic's Claude Sonnet 5, priced at $2/$10 per million tokens during an introductory period (standard $3/$15 after August 2026), matches or nearly matches the flagship Claude Opus 4.8 on knowledge …

21:09
2026-08-11
byteiota.com
artificial-intelligence

Grok 4.6 Is Here: xAI’s Post-Training Bet Against Rivals

XAI released Grok 4.6 on August 7, a language model built on the same 1.5 trillion-parameter V9 foundation as Grok 4.5 but with improved post-training, including better supervised fine-tuning and rein…

04:01
2026-08-04
trae1oung.github.io
artificial-intelligence

SWE-Touch: Benchmarking Coding Agents When Users Touch the Code

Researchers introduced SWE-Touch, a benchmark evaluating coding agents' ability to repair software after users directly edit code, finding that most models' performance drops significantly when user e…

00:00
2026-08-04
mindstudio.ai
artificial-intelligence

Qwen 3.8 Max Explained: Alibaba's 2.4 Trillion Parameter Model

Alibaba released Qwen 3.8 Max, a 2.4 trillion parameter open-weight AI model with 95 billion active parameters, set to become the largest open-weight model once weights are open-sourced about a week a…

15:35
2026-07-29
medley.sh
artificial-intelligence

The harness is the capability multiplier

Medley, a harness system that orchestrates model workers with planning, routing, and acceptance stages, achieved the highest reported results on four public benchmarks: Terminal-Bench 2.1 (93.25% task…

00:00
2026-07-21
vercel.com
artificial-intelligence

Laguna S 2.1 is now available on AI Gateway

Vercel's AI Gateway now offers Laguna S 2.1, an open-weight Mixture-of-Experts model from Poolside that supports up to 1M tokens and agentic coding tasks. The model achieves 70.2% on Terminal-Bench 2.…

09:16
2026-07-13
dev.to
artificial-intelligence

Muse Spark 1.1 + GPT-5.6 launches; Rust 1.97 ships

AI Gateway has become the de facto routing layer for agentic workloads, with Meta's Muse Spark 1.1 and OpenAI's GPT-5.6 family now available through it. Muse Spark 1.1 offers 1M-token context and nati…

14:02
2026-07-10
sourcefeed.dev
artificial-intelligence

The Collapse of SWE-Bench Pro and the Git Scraping Trap

OpenAI retracted its recommendation for Scale AI's SWE-Bench Pro coding benchmark after an audit found roughly 30% of its 731 tasks were broken due to overly strict tests, underspecified prompts, low-…

19:46
2026-07-09
simonwillison.net
artificial-intelligence

The new GPT-5.6 family: Luna, Terra, Sol

OpenAI released the GPT-5.6 family of models in three sizes—Luna, Terra, and Sol—claiming superior long-running agentic performance over Anthropic's Claude Fable 5 on the Agents' Last Exam benchmark, …

02:51
2026-07-09
letsdatascience.com
ai-agents

OpenAI Finds Broken Tasks in SWE-Bench Pro

OpenAI's audit of SWE-Bench Pro found that roughly 30% of the benchmark's tasks are broken, with issues including overly strict hidden tests, underspecified prompts, and low-coverage tests. The findin…

21:03
2026-07-08
openai.com
artificial-intelligence

OpenAI no longer recommends SWE-Bench Pro

OpenAI announced it no longer recommends SWE-Bench Pro as a benchmark for evaluating coding models, citing concerns that the metric may not accurately reflect real-world software engineering performan…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics