Two "Codex CLI" models on the same benchmark: the harness hides the model Specific Labs released Real-SWE, an enterprise-code SWE benchmark whose leaderboard shows that the same CLI harness can produce more than a 2x score gap depending on the underlying model. Results such as Codex CLI scoring 16% versus 34%, and Claude Code reaching 38.8% with Fable 5.1 versus 28.8% with GLM 5.3, indicate that harness names like "Claude Code" or "Codex CLI" describe a routing layer rather than a model. The benchmark frames each result as a model-and-harness combination, and the writeup urges teams evaluating coding agents to pin both the model and the harness behind any reported score. Specific Labs dropped Real-SWE, an enterprise-code SWE benchmark, and the leaderboard is a great study in why you should never read "Claude Code" or "Codex CLI" as a model name. Same harness, two different brains: Same vendor's CLI, same harness, same benchmark. More than a 2x gap. If you'd just read "Codex CLI scored 16%," you'd write off the tool. If you read "Codex CLI scored 34%," you'd maybe believe it. Neither reading is right, because the harness is just the routing layer, not the thing that does the thinking. It's the same on the Claude side. Fable 5.1 on Claude Code is 38.8%. GLM 5.3 on Claude Code is 28.8%. One harness name, ten points apart. Real-SWE is careful about this, actually: they frame every result as a model-and-harness combination, not a model in isolation. That's the honest way to present it, and most vendors won't do it because a low number under their own tooling looks bad. This matters for teams actually shopping for coding agents, because "we use Claude Code" says nothing about the skill of the agent you get. You've picked a route, not a brain. The model swap is the biggest lever, and it's completely invisible in the marketing. The eval habit I'd take from this: when someone hands you a benchmark score, pin both halves. Model. Harness. Conventions, context carry-over, tool loop, judge. If a tool vendor won't tell you which model a number belongs to, that's a red flag, not a detail. The scaffold can move a score more than reasoning effort does.