Giving models the tools they were trained with can make them worse Giving models the exact shell and file-editing tools they were trained with made them less efficient in a MindRoom test: with only Codex's tools, GPT-6 Astra did 84% more work and used 59% more tokens, while Claude Sonnet 5.5 did 15% more work with Claude Code's tools. Only Claude Opus 4.8 worked more efficiently with its native tools, suggesting models generalize beyond their training harness and that what matters is how much the agent resends with every request. The author tested this by matching each model's tool definitions to its native harness — Claude Code's Bash, Read, Edit and Write for Claude, Codex's exec_command and apply_patch for GPT — against MindRoom's own tools. A few nights ago, lying in bed, I had what felt like a galaxy-brain idea. I dictated it into my watch so I would still remember it in the morning. Every frontier model learns to solve coding tasks with tools through reinforcement learning RL , and I assume that training happens inside its vendor’s own coding agent: Anthropic’s models in Claude Code, and OpenAI’s models in Codex. That is a big assumption, since the labs do not publish their training environments, but it makes business sense: they sell those agents, so they have every reason to make their models work best in them. Every other harness, like pi https://github.com/badlogic/pi-mono , opencode https://opencode.ai/ , or my own MindRoom https://www.nijho.lt/post/mindroom/ , gives the model a set of tools of its own design for running commands and editing files. The models are usually smart enough to figure those out, but not always. Sometimes a model assumes an edit tool works like the one it was trained with, and uses it wrong. This summer, that happened to Claude Opus 4.8 in pi, which I come back to below why-i-believed-it . So the idea was simple: in MindRoom, give each model the same shell and file-editing tools as its native harness, and switch them automatically depending on which model is answering. Claude would see Claude Code’s Bash , Read , Edit , and Write . GPT would see Codex’s exec command and apply patch . And because the smallest agents do surprisingly well, I would go one step further: a minimal agent whose only tool is the shell each model knows from its training. I was beyond excited. I could not really share that excitement locally, because my wife did not care, so I am sharing it with the internet instead. Then I measured it against MindRoom’s own tools, and the familiar ones did not make the models more efficient. With only Codex’s tools, GPT-6 Astra did 84% more work, taking more steps and reading and writing more along the way, and used 59% more tokens. With Claude Code’s tools, Claude Sonnet 5.5 did 15% more work, and Claude only came out a few tokens cheaper because the descriptions I wrote for those tools are slightly shorter than MindRoom’s. Only one model clearly worked more efficiently with its own tools: Claude Opus 4.8, the model from the pi story. The models seem to generalize beyond the environment they were trained in, and what still matters is how much the agent sends with every request. “Harness” gets thrown around a lot, but there is no magic in it. Every coding agent talks to the model the same way: each request contains a system prompt, the definition of every tool the model may call, and the conversation so far. When the model answers with a tool call, the harness runs the tool, adds the result to the conversation, and sends everything again. So from the model’s side, a harness is three things: its system prompt, its tool definitions, and what comes back when a tool runs. The first two are just text, so they are easy to copy, and the tool definitions are effectively the harness’s API: change them, and the model sees a different harness. The third is code, which takes more work to copy. Because every request resends everything, a task’s tokens come in two parts. The fixed part is what the agent sends before anything happens: its system prompt, its tool definitions, and the task, which every request resends. The work is everything the steps add on top: the model’s replies, the tool results, and the growing conversation that every later request resends too. When I say below that a model did more work, I mean that second part. Two things convinced me: how well the smallest agents do, and what goes wrong when a model meets unfamiliar tools. Pi made its name by being small. When Mario Zechner introduced pi https://mariozechner.at/posts/2025-11-30-pi-coding-agent/ last November, its system prompt and tool definitions together came in below 1,000 tokens, and it had four tools: read , write , edit , and bash . He also wrote that “pi does not and will not support MCP”, the Model Context Protocol through which most agents load outside tools. And he pointed to Terminus 2, a minimal agent from the Terminal-Bench team, which was “holding its own against agents with far more sophisticated tooling”. In August, DeepSeek released DeepSeek Harness https://github.com/deepseek-ai/deepseek-harness , which was all over Hacker News https://news.ycombinator.com/item?id=49285244 and the local AI subreddits I read. Its standard setup comes with a full set of tools. But for benchmarks, DeepSeek points to its minimal profile https://github.com/deepseek-ai/deepseek-harness/blob/d743267388641bc76f17c45ce8b4c231aed1d32c/BENCHMARK.md . In early September, that profile had exactly one tool, a persistent bash , and a one-line system prompt https://github.com/deepseek-ai/deepseek-harness/blob/63795eaa5cba78cc0a92d95fccd8db5523d7f050/packages/bundle/sdk-minimal/README.md L78-L84 : “You are a helpful software engineer assistant.” Ten days later, it already had a second tool, working directory ; staying that small seems to be hard. That is what inspired the minimal mode I added to MindRoom in September, which gives an agent a short prompt and a single bash tool. The case that got me thinking about this was Claude Opus 4.8 in pi. In July, a pi user reported that about 20% of its edits failed in some sessions https://github.com/earendil-works/pi/issues/6278 . Pi’s edit tool takes a list of replacements, and Opus 4.8 kept adding made-up fields to them, like in file , matchCase , or newText2 . Claude Sonnet 5 did it too, while Opus 4.7 and older Claude models never did in the same tests. Armin Ronacher, one of pi’s maintainers, dug into it https://github.com/earendil-works/pi/issues/6278 issuecomment-4883362982 and compared it with what Claude Code does with the tool calls it receives. It turns out that Claude Code quietly repairs a lot of them: old str and old string . \uXXXX escapes in strings. His hypothesis is that RL itself might be the cause. If a model is trained in a harness that absorbs these mistakes, a slightly wrong tool call still completes the task and still gets rewarded, so nothing teaches the model not to make it. Pi now ignores unknown fields too https://github.com/earendil-works/pi/commit/a1b336d73e13b53949ff629800081185d3e4694e , just like Claude Code. 1 fn:1 Then pi itself strayed from its minimal roots. Pi 1.0 came out on October 1 and added MCP after all https://earendil.com/posts/pi-1-0/ . The Register https://www.theregister.com/ai-and-ml/2026/10/02/pi-coding-agent-pulls-a-180-and-adds-mcp-support/5300678 called it a 180. On Hacker News https://news.ycombinator.com/item?id=49926069 and r/LocalLLaMA https://www.reddit.com/r/LocalLLaMA/comments/1wvffcr/pi 10 released mcp support now included by default/ , people worried that “Pi’s days as a nice minimal agent TUI are numbered” https://news.ycombinator.com/item?id=49926549 . What struck me was how the maintainers explained it. Armin Ronacher https://news.ycombinator.com/item?id=49926377 wrote that the models “are trained on their respective harnesses and we’re not here to fight their behavior”, and Mario Zechner https://news.ycombinator.com/item?id=49926840 that “we follow what the models are trained on”. That is exactly the native-tools half of my plan. My plan combined both: a minimal agent, with each model’s native shell as its only tool. Terminal-Bench also said something I chose to ignore: agents like Terminus 2 do well with tools no model was trained on. I did not believe that part. In opencode and pi, I had repeatedly seen models get the edit tool wrong, so I assumed the benchmarks did not carry over to real work. My belief that a model works best in the harness it was trained in was that strong: I expected the native tools to help, and to help most in a minimal agent with nothing but a shell. MindRoom is my open-source platform for AI agents that live in Matrix chat rooms, where several agents can work together. Its agents get shell and file tools under MindRoom’s own names, like run shell command , read file , and edit file . In this pull request https://github.com/mindroom-ai/mindroom/pull/2766 , I added what I call tool dialects, which show a model those tools in the shape of the coding agent it was trained in. | Model | Sees | Instead of | |---|---|---| | Claude | Bash , BashOutput , KillShell , Read , Edit , and Write , as in Claude Code | run shell command , check shell command , kill shell command , read file , edit file , and write file | | GPT | exec command , write stdin , and a freeform apply patch that takes the patch as plain text, as in Codex, and a kill shell command that takes Codex’s session id | run shell command , check shell command , kill shell command , edit file , and write file | Inside MindRoom, nothing changes: approvals, hooks, and the stored history keep MindRoom’s names, and the translation only happens on the way to and from the model. That gives MindRoom a property I like a lot: a conversation can switch from Claude to GPT halfway through, and each model sees the earlier tool calls in its own shape, as if it had made them in its own harness. Matching the native harness exactly turned out to be a stretch, though. Codex has no tool for reading files, so in the pull request, GPT keeps MindRoom’s read file , and both keep MindRoom’s grep , find files , and ls , which makes each native set a mix. For the main comparison, I removed those extra tools, so each model saw only the tools of its own harness. The tools have the same names and arguments as in Claude Code and Codex, but I wrote short descriptions of my own, because Claude Code’s are not openly licensed. What they return is MindRoom’s output, and the system prompt is MindRoom’s too. So keep in mind for everything below: I tested the names and shapes of the tools, not the whole harness. apply patch is freeform: the model writes the patch as plain text instead of JSON arguments. I ported Codex’s patch parser and the code that applies patches, together with Codex’s own test cases, so a patch from a Codex-trained model applies exactly as it would in Codex. Its description comes from Codex too, which is openly licensed. Edit stays an edit file call for GPT, which has no matching tool. I used 12 terminal tasks from an evaluation suite I built for MindRoom, like writing a report from a CSV file, finding the commit that broke a check, or fixing failing tests.