# Benchmark across Claude, OpenCode, Hermes, and other coding agents on 10 SWE-bench tasks

> Source: <https://www.getreadyforagents.com/news/coding-agent-harness-benchmark-10-task/>
> Published: 2026-09-07 20:05:22+00:00

# Benchmark across Claude, OpenCode, Hermes, and other coding agents on 10 SWE-bench tasks

Carlo Capocasa published a 10-task GLM benchmark comparing Claude, OpenCode, Pi, Zcode, Hermes, and 3code on representative SWE-bench verified tasks. 3code solved 9 of 10 tasks using 5 million tokens, while Pi solved 6 of 10 using fewer tokens than other runners-up. Capocasa noted harness performance varies with token efficiency and task completion rates, cautioning that users should validate results against their own heuristics.

## Topics

## Sources

-   Press   [Read article](https://capocasa.dev/10-task-glm-5-3-harness-bench-claude-opencode-pi-zcode-hermes-and-3code)  

## Go deeper

This intelligence is sourced automatically from public sources across the web and synthesised by the Prefactor AI pipeline. Stories are reviewed before publication.
