08:49
2026-08-17
news.ycombinator.com
artificial-intelligence
Skills x2, 108 Runs: Optimising Efficacy and Token Efficiency
A developer's benchmark of coding agents on 108 bugs found that models memorized answers, raising questions about evaluation efficacy. The Signal Bench, created by darvh, showed high scores but flaggeβ¦