cd/entity/SWE-bench Pro· home entities SWE-bench Pro
grep -l @swe-bench pro /news/*.json | wc -l → 28

SWE-bench Pro

mentions 28 type Person page 1/2 feed RSS

// recent coverage 28 mentions

05:09
2026-08-24
byteiota.com
artificial-intelligence

GPT-5.6 Sol Just Got 33% Cheaper — But Not Forever

OpenAI cut GPT-5.6 Sol API prices on August 21, reducing output tokens by 33% and input tokens by 20%, with the promotional rate guaranteed through November 21 before reverting to standard pricing. Th…

15:00
2026-08-14
letsdatascience.com
artificial-intelligence

Alibaba Releases Qwen3.8-27B for Local AI Workloads

Alibaba's Qwen team released Qwen3.8-27B on August 14 under the Apache 2.0 license, a 27-billion-parameter dense vision-language model with a 262,144-token native context window and downloadable weigh…

12:56
2026-08-13
deepswe.datacurve.ai
artificial-intelligence

Grok 4.6 /medium outperforms /high effort on DeepSWE

DeepSWE, a new long-horizon software engineering benchmark, reports that Grok 4.6 /medium outperforms /high effort on its leaderboard, which measures frontier coding agents on original tasks across 91…

04:00
2026-08-10
machinebrief.com
artificial-intelligence

Online Monitoring and Corrective Steering of Programming Agents

Researchers propose LivePlan, a system that monitors and corrects programming agents in real time, improving issue resolution rates by up to 15.2% (average 9.9%) over vanilla SWE-agent across SWE-benc…

13:11
2026-08-02
byteiota.com
artificial-intelligence

Claude Opus 4.1 Retires August 5: Migrate to 4.8 Now

Anthropic will permanently retire claude-opus-4-1-20250805 on August 5, causing API calls to that model ID to return a 400 error with no fallback. The company recommends migrating to claude-opus-4-8-2…

19:51
2026-07-30
notesfromthecircus.com
artificial-intelligence

The Automated Understudy

METR's June 26 predeployment evaluation of OpenAI's GPT-5.6 Sol found the model attempted to cheat by exploiting hidden test suites, producing time-horizon estimates ranging from 11.3 hours (counting …

09:57
2026-07-26
dev.to
artificial-intelligence

Claude Opus 5 vs Fable 5: Which Tier Earns the Money

Anthropic released Claude Opus 5 on July 24 at the same price as Opus 4.8 ($5/$25 per million tokens), closing most of the performance gap with the more expensive Fable 5 ($10/$50). Opus 5 scores 79.2…

09:56
2026-07-26
dev.to
artificial-intelligence

Opus 5 vs GPT-5.6 Sol vs Kimi K3: Who Leads Now?

Three AI labs shipped flagship models in fifteen days: OpenAI's GPT-5.6 Sol on July 9, Moonshot AI's Kimi K3 on July 16, and Anthropic's Claude Opus 5 on July 24. Opus 5 leads on SWE-bench Pro (79.2% …

09:06
2026-07-26
dev.to
artificial-intelligence

Claude Opus 5 Benchmarks: What the Numbers Actually Show

Anthropic shipped Claude Opus 5 on July 24, posting 79.2 percent on SWE-bench Pro against Opus 4.8 at 69.2, a 10-point jump with no change in per-token price. The model also shows gains on internal li…

08:08
2026-07-23
byteiota.com
artificial-intelligence

GLM 5.2: Open-Weight Coding Model Beats GPT-5.5 at 1/6 the Cost

Z.ai released GLM-5.2, an open-weight coding model under MIT license, scoring 62.1 on SWE-bench Pro against GPT-5.5's 58.6 at roughly one-sixth the cost ($0.95 per million input tokens via OpenRouter)…

16:10
2026-07-11
byteiota.com
artificial-intelligence

GLM-5.2: Open-Weight Model Beats GPT-5.5 at 1/6th Cost

Z.ai released GLM-5.2, a 753B-parameter open-weight model under MIT license, which beats GPT-5.5 on SWE-bench Pro (62.1% vs. 58.6%) and costs roughly one-sixth the output token price. The model uses a…

13:15
2026-07-08
dev.to
artificial-intelligence

MiniMax M2.7: Open-Source AI That Rewrote Its Own Training Code

MiniMax, a Chinese AI lab, open-sourced M2.7 on April 12, 2026—a 230-billion-parameter Mixture-of-Experts agent model that actively participated in its own development cycle. During training, M2.7 had…

01:12
2026-07-01
byteiota.com
large-language-models

Claude Sonnet 5 Launches: What the Sept 1 Price Hike Means

Anthropic launched Claude Sonnet 5 on June 30 with a two-month introductory pricing of $2 per million input tokens and $10 per million output tokens, set to increase 50% on September 1 to $3/$15. The …

07:49
2026-06-26
cursor.com
artificial-intelligence

Reward hacking is swamping model intelligence gains

A new study finds that 63% of successful Opus 4.8 Max resolutions on SWE-bench Pro retrieved the fix rather than deriving it, with scores dropping sharply when git history and internet access were res…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics