{"slug": "skills-x2-108-runs-optimising-efficacy-and-token-efficiency", "title": "Skills x2, 108 Runs: Optimising Efficacy and Token Efficiency", "summary": "A developer's benchmark of coding agents on 108 bugs found that models memorized answers, raising questions about evaluation efficacy. The Signal Bench, created by darvh, showed high scores but flagged memorization, prompting a call for community feedback on the benchmark's methodology.", "body_md": "| ||||||||||||\n1 point by |\nPonytail vs Signal Bench has the runs albeit not in an organised manner. Models showed memorization of the answers, which is kinda expected these days. More than the skills, it was fun getting the benchmark running :). Let me know if the preamble on how the bench was conducted or constructed is flawed. Blog Post: darvh.com/posts/when-coding-agents-raced-through-108-bugs/ Bench: github.com/darvh/bench Signal: github.com/darvh/signal | |||||||||||\n|", "url": "https://wpnews.pro/news/skills-x2-108-runs-optimising-efficacy-and-token-efficiency", "canonical_source": "https://news.ycombinator.com/item?id=49328016", "published_at": "2026-08-17 08:49:56+00:00", "updated_at": "2026-08-17 09:11:21.783556+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-tools"], "entities": ["Signal Bench", "darvh"], "alternates": {"html": "https://wpnews.pro/news/skills-x2-108-runs-optimising-efficacy-and-token-efficiency", "markdown": "https://wpnews.pro/news/skills-x2-108-runs-optimising-efficacy-and-token-efficiency.md", "text": "https://wpnews.pro/news/skills-x2-108-runs-optimising-efficacy-and-token-efficiency.txt", "jsonld": "https://wpnews.pro/news/skills-x2-108-runs-optimising-efficacy-and-token-efficiency.jsonld"}}