cd /news/ai-tools/10-task-glm-5-3-harness-bench-claude… · home topics ai-tools article
[ARTICLE · art-122198] src=capocasa.dev ↗ pub= topic=ai-tools verified=true sentiment=· neutral

10-task GLM 5.3 harness bench: claude, opencode, pi, zcode, hermes and 3code

In a 10-task SWE-bench verified harness benchmark, 3code solved 9 of 10 tasks using 5 million tokens, while Claude Code, opencode, pi, zcode, and hermes were also tested. The benchmark, conducted by independent developer Carlo from Munich, found 3code the most token-efficient among top solvers, with opencode also solving 9 tasks but using twice the tokens. The results are intended as a rough guide, with the author noting that token usage and cache rates vary significantly across harnesses.

read3 min views1 publishedSep 6, 2026
10-task GLM 5.3 harness bench: claude, opencode, pi, zcode, hermes and 3code
Image: Capocasa (auto-discovered)

I'm performing a series of harness benchmarks on the same 10 SWE-bench verified tasks representatively chosen for difficulty. This is far from a perfect measure and unpublished model improvements mean that results are far from deterministic, but it does allow for a rough evaluation without breaking the bank. As with all benchmarks- use your own heuristics to confirm! Use the harness and check your feel- do you get a sense you getting more mileage? That's what couts because it's what you can actually measure.

Having said that, this sort of benchmarks have been positively great for spotting harness blunders.

Note that none of these results, even when taken with a grain of salt, mean you should necessarily chose one harness over another. If you have an unlimited budget and just like Claude Code, by all means go for it. But if you want absolute maximum savings and you like very terse command line agents- give 3code a try!

So here we go.

In the graphic there are 5 columns- harness name (the latest at time of writing was chosen), amount of solved tasks of the 10, an indication of used tokens as well as total tokens, cache rate, and percentage of tokens used relative to the winner.

The picture is pretty drastic- 3code used 5 million tokens and solved 9 of the 10 tasks. The runner up in token usage is pi, which only solved 6 of the 10 tasks in the alotted time- the runner up in completed tasks is opencode, the only other harness to solve 9 of 10 tasks ('astropi' is the hard one that can be flipped).

zcode was hyper efficient on the tasks it was successful at, but dug its heels in on the tasks it was unsuccessful at and used a lot of tokens before timing out. Hermes was expensive and did okay- I suppose that's a pretty good result since it is not strictly a coding model and more of an agentic model where cheaper models to moderate difficulty coding to expand capabilities.

As expected, Claude Code positively splurges tokens, consistent with Anthropics hyper maximalist approach to all aspects of performance. That it didn't get astropi might have been bad luck.

opencode's excellent cache rate is notable- however, given it uses twice the tokens, it's still the more expensive choice. Personally, in most of my projects, cached tokens are the largest expense so I believe in reducing overall token usage more than in increasing cache rate. After all, if you have more tokens used again and again, there is a greater opportunity to cache but then it still costs. Having said that- this test certainly informs 3code development towards finding unnecessarily uncached tokens and 99.2% on an independent benchmark is very impressive. Our lowly 95% numbers might still be good enough though!

pi continues to impress how far it can go without having model specific tuning, but I do feel that the 3code approach of having a model family specific system prompt and also finetuning was shown to be paying off. I think what's happening is that this run's 3code version benefitted from stealing all the zcode GLM 5.3 specific finetuning, but then not having the overly aggressive task completion policy.

I intend to continue to create harness benchmarks and am working on funding mor extensive tests, and also different kinds of tests- cost is the limiting factor right now since I'm independent and not associated with any provider.

The tokens for this test were provided by Zhipu's startup program, but they didn't influence the results in any way- they gave me the tokens and cut me loose. I also converged on GLM as the main 3code dev tool first, and asked them to give me tokens second. But I'm still interested in doing tests with other companies as well, Zhipu was just the first to give a positive response to my requests. Thank you!

Servus aus München, Carlo

── more in #ai-tools 4 stories · sorted by recency
── more on @3code 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/10-task-glm-5-3-harn…] indexed:0 read:3min 2026-09-06 ·