22:00
2026-09-06
capocasa.dev
ai-tools
10-task GLM 5.3 harness bench: claude, opencode, pi, zcode, hermes and 3code
In a 10-task SWE-bench verified harness benchmark, 3code solved 9 of 10 tasks using 5 million tokens, while Claude Code, opencode, pi, zcode, and hermes were also tested. The benchmark, conducted by iā¦