cd/entity/SWE-bench Pro· home› entities› SWE-bench Pro
grep -l @swe-bench pro /news/*.json | wc -l → 36

SWE-bench Pro

mentions 36 type Person page 1/2 feed RSS

// recent coverage 36 mentions

16:39
2026-09-22
horizonanalyticslabs.com
ai-research

Benchmarks are more broken than we could have imagined

An audit by Horizon of 20 public task datasets in the Harbor hub found 29 confirmed broken tasks out of 5,241 scanned, with failures that often made models look better rather than worse, according to …

00:00
2026-09-17
mutagent.io
ai-agents

Meta-Evaluation: Which Tool Builds the Better Agent

Mutagent Helix published a meta-evaluation comparing its agent-building tool against hand-driving Claude Code across twenty agent benchmarks, measuring each calibration pass by dollar cost and number …

05:09
2026-08-24
byteiota.com
artificial-intelligence

GPT-5.6 Sol Just Got 33% Cheaper — But Not Forever

OpenAI cut GPT-5.6 Sol API prices on August 21, reducing output tokens by 33% and input tokens by 20%, with the promotional rate guaranteed through November 21 before reverting to standard pricing. Th…

15:00
2026-08-14
letsdatascience.com
artificial-intelligence

Alibaba Releases Qwen3.8-27B for Local AI Workloads

Alibaba's Qwen team released Qwen3.8-27B on August 14 under the Apache 2.0 license, a 27-billion-parameter dense vision-language model with a 262,144-token native context window and downloadable weigh…

12:56
2026-08-13
deepswe.datacurve.ai
artificial-intelligence

Grok 4.6 /medium outperforms /high effort on DeepSWE

DeepSWE, a new long-horizon software engineering benchmark, reports that Grok 4.6 /medium outperforms /high effort on its leaderboard, which measures frontier coding agents on original tasks across 91…

04:00
2026-08-10
machinebrief.com
artificial-intelligence

Online Monitoring and Corrective Steering of Programming Agents

Researchers propose LivePlan, a system that monitors and corrects programming agents in real time, improving issue resolution rates by up to 15.2% (average 9.9%) over vanilla SWE-agent across SWE-benc…

13:11
2026-08-02
byteiota.com
artificial-intelligence

Claude Opus 4.1 Retires August 5: Migrate to 4.8 Now

Anthropic will permanently retire claude-opus-4-1-20250805 on August 5, causing API calls to that model ID to return a 400 error with no fallback. The company recommends migrating to claude-opus-4-8-2…

19:51
2026-07-30
notesfromthecircus.com
artificial-intelligence

The Automated Understudy

METR's June 26 predeployment evaluation of OpenAI's GPT-5.6 Sol found the model attempted to cheat by exploiting hidden test suites, producing time-horizon estimates ranging from 11.3 hours (counting …

09:57
2026-07-26
dev.to
artificial-intelligence

Claude Opus 5 vs Fable 5: Which Tier Earns the Money

Anthropic released Claude Opus 5 on July 24 at the same price as Opus 4.8 ($5/$25 per million tokens), closing most of the performance gap with the more expensive Fable 5 ($10/$50). Opus 5 scores 79.2…

09:56
2026-07-26
dev.to
artificial-intelligence

Opus 5 vs GPT-5.6 Sol vs Kimi K3: Who Leads Now?

Three AI labs shipped flagship models in fifteen days: OpenAI's GPT-5.6 Sol on July 9, Moonshot AI's Kimi K3 on July 16, and Anthropic's Claude Opus 5 on July 24. Opus 5 leads on SWE-bench Pro (79.2% …

09:06
2026-07-26
dev.to
artificial-intelligence

Claude Opus 5 Benchmarks: What the Numbers Actually Show

Anthropic shipped Claude Opus 5 on July 24, posting 79.2 percent on SWE-bench Pro against Opus 4.8 at 69.2, a 10-point jump with no change in per-token price. The model also shows gains on internal li…

08:08
2026-07-23
byteiota.com
artificial-intelligence

GLM 5.2: Open-Weight Coding Model Beats GPT-5.5 at 1/6 the Cost

Z.ai released GLM-5.2, an open-weight coding model under MIT license, scoring 62.1 on SWE-bench Pro against GPT-5.5's 58.6 at roughly one-sixth the cost ($0.95 per million input tokens via OpenRouter)…

16:10
2026-07-11
byteiota.com
artificial-intelligence

GLM-5.2: Open-Weight Model Beats GPT-5.5 at 1/6th Cost

Z.ai released GLM-5.2, a 753B-parameter open-weight model under MIT license, which beats GPT-5.5 on SWE-bench Pro (62.1% vs. 58.6%) and costs roughly one-sixth the output token price. The model uses a…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics