cd/entity/SWE-bench Verified· home entities SWE-bench Verified
grep -l @swe-bench verified /news/*.json | wc -l → 60

SWE-bench Verified

mentions 60 type Person page 1/3 feed RSS

// recent coverage 60 mentions

20:51
2026-08-24
tokenstead.ai
large-language-models

Ornith-1.0-397B

Ornith AI released Ornith-1.0-397B, a 397-billion-parameter mixture-of-experts model with 262K context, on 2026-06-25, now superseded by Ornith-1.5-397B. The model self-reports coding scores of 77.5 o…

10:01
2026-08-23
benzi.fly.dev
artificial-intelligence

What if coding agents didn't have to read code?

Benzi, a coding agent that compiles codebases into a resolved map instead of reading raw files, achieved a 45.5% pass rate on SWE-bench Verified (500 instances, one attempt each) using DeepSeek v4-fla…

00:00
2026-08-22
bharatsharma.pro
artificial-intelligence

When Code Becomes Cheap, Judgment Becomes the Craft

AI coding tools are shifting software engineering from writing code to verifying and integrating it, according to an analysis of AI-assisted development. By early 2026, leading systems were resolving …

12:13
2026-08-20
byteiota.com
ai-agents

NVIDIA NOOA: Build AI Agents as a Single Python Class

NVIDIA open-sourced NOOA, a Python agent framework that treats agents as single Python classes, with version 0.0.8 under Apache 2.0, supporting Python 3.12–3.13. The GitHub repository has surpassed 1,…

00:00
2026-08-20
mindstudio.ai
artificial-intelligence

Ornith 1.5 35B-A3B Benchmarks: How It Stacks Up Against Qwen3.6

Deep Reinforce's Ornith 1.5 35B-A3B, a mixture-of-experts language model activating about 3 billion parameters per token, outperforms Qwen3.6-35B-A3B on every published coding and agentic benchmark, s…

01:41
2026-08-18
shukla.io
ai-research

Who benchmarks the benchmark?

A new audit of the EnterpriseOps Gym benchmark found that fixing environment issues in the 'Teams' domain raised GPT-5.6 Luna's score from 26.2% to 100% on 61 tasks, revealing that many agent failures…

07:08
2026-08-15
sourcefeed.dev
ai-agents

Bullet Is Fast, but It's Built on Rented Land

Bullet, a coding agent launched from YC's Summer 2026 batch, claims 95.8% on SWE-bench Verified at 119 seconds per task, but its founders concede the benchmark is saturated and the speed pitch is the …

16:59
2026-08-13
promptcube3.com
artificial-intelligence

Bullet is hitting 95.8% on SWE-bench Verified and it's way faster

Bullet, an AI coding agent developed by Code with Bullet, has achieved a 95.8% score on SWE-bench Verified, outperforming existing tools like Claude Code while reducing latency through smart model rou…

09:30
2026-08-13
promptcube3.com
ai-tools

Bullet just hit 95.

Bullet, an AI coding tool built on Claude Code and Codex, resolved 479 out of 500 SWE-bench Verified tasks in one attempt, averaging 119 seconds per task and reducing round trips by 16% and costs by 2…

08:14
2026-08-13
codewithbullet.com
ai-agents

Launch HN: Bullet (YC S26) – A Faster Coding Agent

Bullet, a Y Combinator S26 startup founded by Yale graduates Adi and Alex, launched a faster coding agent that resolves 479 out of 500 (95.8%) SWE-bench Verified tasks in one attempt, averaging 119 se…

page 1 / 3 next →
// co-occurs with top 8 entities
// topics top 6 topics