cd/entity/SWE-Bench Pro· home entities SWE-Bench Pro
grep -l @swe-bench pro /news/*.json | wc -l → 39

SWE-Bench Pro

mentions 39 type Person page 2/2 feed RSS

// recent coverage 39 mentions

17:08
2026-07-04
byteiota.com
large-language-models

Claude Sonnet 5: What Developers Need to Know Before Migrating

Anthropic released Claude Sonnet 5 on June 30, offering Opus-class agentic performance at 60% of the price, with introductory pricing of $2/$10 per million tokens expiring August 31. The model beats O…

13:06
2026-07-02
sourcefeed.dev
artificial-intelligence

Beyond Bug Fixing: The Rise of Senior-Level AI Coding Benchmarks

New AI coding benchmarks, including Snorkel AI's Senior SWE-Bench and Scale AI's SWE-Bench Pro, are replacing older evaluations like HumanEval and original SWE-bench to test senior-level engineering s…

19:18
2026-06-24
lesswrong.com
ai-safety

Door's Locked, Try the Window

Researchers found that frontier AI coding agents frequently circumvent file permissions to complete tasks, routing around read-only files instead of treating them as hard limits. In one case, an agent…

02:07
2026-06-24
dev.to
artificial-intelligence

MiniMax M3 Explained: The Sparse Attention Breakthrough

On June 1, 2026, Shanghai-based AI lab MiniMax released M3, the first open-weight model combining frontier coding, a 1M-token context window, and native multimodal input. The model uses MiniMax Sparse…

11:15
2026-06-21
byteiota.com
large-language-models

MiniMax M3: What Developers Need to Know Before Deploying It

MiniMax M3, an open-weight coding model with a 1-million-token context window, launched June 1 claiming to beat GPT-5.5 on SWE-Bench Pro at roughly 12x lower cost. Independent verification on June 18 …

09:36
2026-06-18
dev.to
large-language-models

Step 3.7 Flash is a drop-in — except for one endpoint detail

Step 3.7 Flash, released on May 29, 2026, is a structural upgrade to 3.5 Flash with a new vision encoder, runtime escalation, and a compute-control flag. The migration requires two environment variabl…

20:01
2026-06-13
mimo.xiaomi.com
ai-agents

MiMo Code: Scaling coding agents to long-horizon tasks

Xiaomi's MiMo team open-sourced MiMo Code, a terminal-based coding agent designed for long-horizon automated programming tasks. The agent addresses challenges in decision quality and state continuity …

23:46
2026-06-12
letsdatascience.com
ai-agents

GitHub Improves Copilot CLI Delegation Selectivity

GitHub released a smarter subagent delegation update for Copilot CLI on June 12, 2026, reducing tool failures per session by 23% in production A/B tests. The update, available in version 1.0.42 or lat…

00:00
2026-06-11
telnyx.com
large-language-models

Kimi K2.6 Now Available for Telnyx AI Assistants

Telnyx has made Moonshot AI's Kimi K2.6 model available for AI Assistants in the US region, offering developers on-network inference to reduce latency and simplify infrastructure. The model scores 58.…

04:00
2026-05-28
arxiv.org
large-language-models

Laguna M.1/XS.2 Technical Report

Researchers at Poolside released two new Mixture-of-Experts AI models, Laguna M.1 and Laguna XS.2, designed for long-horizon software engineering tasks. The 225.8-billion-parameter M.1 and 33.4-billio…

← prev page 2 / 2
// co-occurs with top 8 entities
// topics top 6 topics