cd/entity/SWE-bench Pro· home entities SWE-bench Pro
grep -l @swe-bench pro /news/*.json | wc -l → 28

SWE-bench Pro

mentions 28 type Person page 2/2 feed RSS

// recent coverage 28 mentions

12:09
2026-06-25
byteiota.com
large-language-models

GLM-5.2 Beats GPT-5.5 at Coding for One-Sixth the Price

Z.AI's open-weight GLM-5.2 model outperforms GPT-5.5 on the SWE-bench Pro coding benchmark, scoring 62.1 versus 58.6, while costing $1.40 per million input tokens compared to GPT-5.5's $8.00. Released…

09:21
2026-06-25
gopeekapp.blogspot.com
large-language-models

Claude Opus 4.5 vs. GLM-5.2

Anthropic's Claude Opus 4.5 and Zhipu AI's GLM-5.2 are frontier reasoning models competing on cost, context length, and coding benchmarks. GLM-5.2 offers a 1M-token context window and leads by 20.3 po…

11:10
2026-06-24
byteiota.com
ai-research

SWE-bench Pro: How to Read the Coding Agent Leaderboard

OpenAI abandoned SWE-bench Verified on February 23, 2026, after finding 59.4% of its hardest failed tests were broken and training data contamination inflated scores. Its replacement, SWE-bench Pro fr…

09:01
2026-06-17
huggingface.co
large-language-models

GLM-5.2: Built for Long-Horizon Tasks

Zhipu AI released GLM-5.2, an open-source large language model with a 1M-token context designed for long-horizon coding tasks. The model introduces IndexShare to reduce computational costs and achieve…

23:20
2026-06-11
cryptobriefing.com
artificial-intelligence

Xiaomi’s MiMo Code outperforms Claude Code in 200+ step tasks

Xiaomi released MiMo Code V0.1, an open-source AI coding agent, on June 11 under an MIT license, achieving an 86.7% score on Terminal-Bench 2.0 compared to Claude Code's 65.4% on tasks exceeding 200 s…

14:10
2026-06-02
lesswrong.com
large-language-models

Claude Opus 4.8: Capabilities and Reactions

Anthropic released Claude Opus 4.8 on Tuesday, pricing the new model at $5 per million input tokens and $25 per million output tokens — the same as its predecessor. The company claims the model is its…

← prev page 2 / 2
// co-occurs with top 8 entities
// topics top 6 topics