Carlo Capocasa published a 10-task GLM benchmark comparing Claude, OpenCode, Pi, Zcode, Hermes, and 3code on representative SWE-bench verified tasks. 3code solved 9 of 10 tasks using 5 million tokens, while Pi solved 6 of 10 using fewer tokens than other runners-up. Capocasa noted harness performance varies with token efficiency and task completion rates, cautioning that users should validate results against their own heuristics.
Topics #
Sources #
- Press Read article
Go deeper #
This intelligence is sourced automatically from public sources across the web and synthesised by the Prefactor AI pipeline. Stories are reviewed before publication.