Benchmark across Claude, OpenCode, Hermes, and other coding agents on 10 SWE-bench tasks Carlo Capocasa published a 10-task benchmark comparing coding agents Claude, OpenCode, Pi, Zcode, Hermes, and 3code on SWE-bench verified tasks, with 3code solving 9 of 10 tasks using 5 million tokens and Pi solving 6 of 10 with fewer tokens. Capocasa cautioned that harness performance varies with token efficiency and task completion rates, advising users to validate results against their own heuristics. Benchmark across Claude, OpenCode, Hermes, and other coding agents on 10 SWE-bench tasks Carlo Capocasa published a 10-task GLM benchmark comparing Claude, OpenCode, Pi, Zcode, Hermes, and 3code on representative SWE-bench verified tasks. 3code solved 9 of 10 tasks using 5 million tokens, while Pi solved 6 of 10 using fewer tokens than other runners-up. Capocasa noted harness performance varies with token efficiency and task completion rates, cautioning that users should validate results against their own heuristics. Topics Sources - Press Read article https://capocasa.dev/10-task-glm-5-3-harness-bench-claude-opencode-pi-zcode-hermes-and-3code Go deeper This intelligence is sourced automatically from public sources across the web and synthesised by the Prefactor AI pipeline. Stories are reviewed before publication.