cd /news/ai-research/two-codex-cli-models-on-the-same-ben… · home topics ai-research article
[ARTICLE · art-128610] src=dev.to ↗ pub= topic=ai-research verified=true sentiment=· neutral

Two "Codex CLI" models on the same benchmark: the harness hides the model

Specific Labs released Real-SWE, an enterprise-code SWE benchmark whose leaderboard shows that the same CLI harness can produce more than a 2x score gap depending on the underlying model. Results such as Codex CLI scoring 16% versus 34%, and Claude Code reaching 38.8% with Fable 5.1 versus 28.8% with GLM 5.3, indicate that harness names like "Claude Code" or "Codex CLI" describe a routing layer rather than a model. The benchmark frames each result as a model-and-harness combination, and the writeup urges teams evaluating coding agents to pin both the model and the harness behind any reported score.

by read1 min views3 publishedSep 14, 2026

Specific Labs dropped Real-SWE, an enterprise-code SWE benchmark, and the leaderboard is a great study in why you should never read "Claude Code" or "Codex CLI" as a model name.

Same harness, two different brains:

Same vendor's CLI, same harness, same benchmark. More than a 2x gap. If you'd just read "Codex CLI scored 16%," you'd write off the tool. If you read "Codex CLI scored 34%," you'd maybe believe it. Neither reading is right, because the harness is just the routing layer, not the thing that does the thinking.

It's the same on the Claude side. Fable 5.1 on Claude Code is 38.8%. GLM 5.3 on Claude Code is 28.8%. One harness name, ten points apart. Real-SWE is careful about this, actually: they frame every result as a model-and-harness combination, not a model in isolation. That's the honest way to present it, and most vendors won't do it because a low number under their own tooling looks bad.

This matters for teams actually shopping for coding agents, because "we use Claude Code" says nothing about the skill of the agent you get. You've picked a route, not a brain. The model swap is the biggest lever, and it's completely invisible in the marketing.

The eval habit I'd take from this: when someone hands you a benchmark score, pin both halves. Model. Harness. Conventions, context carry-over, tool loop, judge. If a tool vendor won't tell you which model a number belongs to, that's a red flag, not a detail. The scaffold can move a score more than reasoning effort does.

── more in #ai-research 4 stories · sorted by recency
── more on @specific labs 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/two-codex-cli-models…] indexed:0 read:1min 2026-09-14 ·