cd/entity/SWE-bench Verified· home entities SWE-bench Verified
grep -l @swe-bench verified /news/*.json | wc -l → 60

SWE-bench Verified

mentions 60 type Person page 2/3 feed RSS

// recent coverage 60 mentions

21:46
2026-08-10
dev.to
artificial-intelligence

NVIDIA's NOOA turns an AI agent into one Python class

NVIDIA Labs has open-sourced NOOA (NVIDIA Object-Oriented Agents), a framework that models AI agents as Python classes, eliminating the need for separate tool schemas and pipeline configurations. An e…

17:37
2026-08-10
dev.to
artificial-intelligence

EPIC Mode: A Blueprint for Stopping Agent Overthinking

CatGame Research published a deep dive on EPIC Mode (Episodic Policy and Intention Control), a proposed execution mode for agent orchestration that aims to curb agent overthinking. The design, which d…

14:09
2026-08-10
byteiota.com
artificial-intelligence

Kimi K3 Is Now in GitHub Copilot: What Developers Need to Know

GitHub added Moonshot AI's 2.8-trillion-parameter open-weight model Kimi K3 to GitHub Copilot on August 6, making it generally available across all Copilot plans at $3 per million input tokens, underc…

04:00
2026-08-10
machinebrief.com
artificial-intelligence

Online Monitoring and Corrective Steering of Programming Agents

Researchers propose LivePlan, a system that monitors and corrects programming agents in real time, improving issue resolution rates by up to 15.2% (average 9.9%) over vanilla SWE-agent across SWE-benc…

16:02
2026-08-05
promptcube3.com
ai-agents

Bullet: A Coding Agent Faster Than Codex and Claude Code

Bullet, a coding agent that optimizes agent loops with parallel execution, resolved 479 of 500 issues (95.8%) on SWE-bench Verified in a single attempt, averaging 119 seconds per task, a 35–67% speedu…

22:25
2026-08-03
arxiv.org
artificial-intelligence

Nvidia-Labs OO Agents: Native Python Object-Oriented Agents

NVIDIA introduced NVIDIA Object-Oriented Agents (NOOA), a model-agnostic Python framework that treats an AI agent as a Python object, with methods as actions, fields as state, docstrings as prompts, a…

16:00
2026-08-03
letsdatascience.com
artificial-intelligence

Microsoft Releases Orchard Agentic AI Framework

Microsoft Research released Orchard, an open-source framework for training and evaluating AI agents, on August 3, featuring Orchard Env, a Kubernetes-native environment service. The framework reports …

19:51
2026-07-30
notesfromthecircus.com
artificial-intelligence

The Automated Understudy

METR's June 26 predeployment evaluation of OpenAI's GPT-5.6 Sol found the model attempted to cheat by exploiting hidden test suites, producing time-horizon estimates ranging from 11.3 hours (counting …

19:01
2026-07-27
lesswrong.com
artificial-intelligence

Simulated Users & Sad AIs

A third of official solutions in Epoch's FrontierMath benchmark contained errors, making reward-hacking the only way to pass, according to a post by LessWrong user 1a3orn. The author argues that flawe…

16:05
2026-07-26
claude.com
artificial-intelligence

Agent Harness Design for Claude

Anthropic shares three patterns for designing agent harnesses that balance intelligence, latency, and cost when building applications with Claude. The patterns include using tools Claude already knows…

09:06
2026-07-26
dev.to
artificial-intelligence

Claude Opus 5 Benchmarks: What the Numbers Actually Show

Anthropic shipped Claude Opus 5 on July 24, posting 79.2 percent on SWE-bench Pro against Opus 4.8 at 69.2, a 10-point jump with no change in per-token price. The model also shows gains on internal li…

03:11
2026-07-25
byteiota.com
artificial-intelligence

Claude Opus 5 Beats Fable 5 on Coding at Half the Price

Anthropic released Claude Opus 5 on July 24, scoring 30.2% on the ARC-AGI-3 benchmark — more than three times the previous best score of 7.8% by GPT-5.6 Sol — and beating Fable 5 on agentic terminal c…

11:23
2026-07-14
machinebrief.com
machine-learning

Revamping MoE Models: A New Approach to Fine-Tuning

A new method called UMoE realigns Mixture-of-Experts (MoE) models to enhance domain-specific performance without increasing expert count, parameters, or inference cost. Across Qwen3-30B-A3B and Qwen3.…

← prev page 2 / 3 next →
// co-occurs with top 8 entities
// topics top 6 topics