13:51
2026-08-15
arxiv.org
artificial-intelligence
QuoteBench: Matched Scores Can Hide Command-Path Failures
A new benchmark, QuoteBench, reveals that matched execution scores for LLM coding agents can hide command-path failures, with replaying the same reply through an added parser lowering success by 55.4 โฆ