cd/entity/ProgramBench· home› entities› ProgramBench
grep -l @programbench /news/*.json | wc -l → 3

ProgramBench

mentions 3 type Organization feed RSS

// recent coverage 3 mentions

21:25
2026-08-27
factory.com
artificial-intelligence

Why coding agents stop early on long-horizon software tasks

A study by the coding agent company Droid found that single-agent runs stop early on long-horizon software tasks because they validate work locally without a complete standard of completion. In a test…

15:29
2026-08-20
promptcube3.com
artificial-intelligence

ProgramBench reverse-engineering benchmark puts LLMs to the test

ProgramBench, a new reverse-engineering benchmark, compiles and fuzzes LLM-generated code against original binaries, with GPT-4o achieving a 34% pass@1 rate, Claude 3.5 Sonnet 28%, and DeepSeek-Coder-…

23:24
2026-08-14
lesswrong.com
artificial-intelligence

Your Agents Are Not Time Aware

Coding agents Claude Code and Codex consistently over-predict their own wall-clock runtime, according to research conducted as part of MATS 10 with Maksym Andriushchenko. The study introduced AgentTim…

// co-occurs with top 8 entities
// topics top 6 topics