{"slug": "claude-code-and-codex-break-on-different-mcp-features", "title": "Claude Code and Codex break on different MCP features", "summary": "A test of ten MCP features against Claude Code 2.1.287 (claude-sonnet-5-5) and Codex 0.162.0 (gpt-5.6-sol) on 9 October 2026 found that five features worked on one agent and failed on the other, with no feature failing on both, according to the Harness Quirks Matrix built on the open-source M3 testing tool. Claude Code dropped tools with an input property named `ids[]`, lost results carrying both text and `structuredContent`, and truncated tool errors to roughly the first and last 5,000 characters, while Codex failed to call a new tool added after a tool-list change and dropped required nested fields when compacting a 14 KB input schema. Eight further failures found by direct binary probes, including Codex dropping tools past 2,048 on a server with 2,049 tools, have not yet run through M3 and are treated as leads.", "body_md": "# Claude Code and Codex break on different MCP features\n\nWe ran Claude Code and Codex against the same ten MCP servers at pinned versions. Five features split them, and they never failed the same one.\n\n- Published\n- Reading time\n- 5 min\n\nWe ran Claude Code and Codex against the same ten MCP servers. Five of the ten features worked on one agent and failed on the other, and the two agents never failed on the same feature.\n\nIf you ship an MCP server, you have probably seen the issue: \"tool doesn't show up in Claude Code\" or \"Codex sends the wrong arguments.\" When we started, about 28 issues like these were open across agent repositories. None of them says which agent, at which version, breaks on which MCP feature.\n\nSo we built a table that does. We call it the Harness Quirks Matrix, and it runs on [M3](https://github.com/sineframe/m3), our open-source tool for [testing MCP servers](https://m3.sineframe.com/mcp-testing) and the agents that use them.\n\n## The results: 5 of 10 MCP features split the agents\n\nOn 9 October 2026 we ran ten MCP features against Claude Code 2.1.287 (`claude-sonnet-5-5`) and Codex 0.162.0 (` gpt-5.6-sol`), three trials each. Claude Code failed three of the features and Codex failed the other two.\n\n| MCP feature | Claude Code 2.1.287 | Codex 0.162.0 | \n|---|---|---|\n| Tool argument named `ids[]` | dropped0/3 | works3/3 | \n| Marker in the middle of a 12,000-character tool error | result lost0/3 | works3/3 | \n| Result with both text and `structuredContent` | result lost0/3 | works3/3 | \n| New tool added after a tool-list change | works3/3 | dropped0/3 | \n| Required nested fields in a 14 KB input schema | works3/3 | dropped0/3 | \n| Root-level `anyOf` in`inputSchema` | works3/3 | works3/3 | \n| `outputSchema` declaring draft-07 | works3/3 | works3/3 | \n| Omitting optional parameters that have defaults | works3/3 | works3/3 | \n| Tool whose full name is 65 characters | works3/3 | works3/3 | \n| `resource_link` block in a result | works3/3 | works3/3 | \n\nHere is what the five failures mean if you maintain an MCP server.\n\nClaude Code never calls a tool whose input property is named `ids[]`. It shows no error, and the tool simply isn't there. OpenAPI-generated servers often produce names like this.\n\nClaude Code also keeps only about the first and last 5,000 characters of a tool error, so anything in the middle never reaches the model. Put the code, the reason and the next step at the start of an error message. And when a result carries both `content` and `structuredContent`, the model reported the structured value and lost the text one.\n\nCodex kept using its old catalog after the server announced a new tool, so it never called the new one. It also compacts large input schemas to fit a budget, and in our 14 KB schema the required nested fields went with them. Raising Codex's per-server schema budget to 20,000 restored the call.\n\nEach result covers one client and model pair at pinned versions, with explicit prompts. A failed cell tells you the data didn't get from the server to the model's answer; it doesn't tell you which layer dropped it. Every cell has its run manifest, binary SHA-256 and wire trace stored with it.\n\n## Eight more failures from direct probes\n\nWe also probed the Claude Code and Codex binaries directly, without M3, and found eight more features that fail on at least one of them. None has run through M3 yet, so treat them as leads until they do.\n\n| MCP feature | Claude Code 2.1.295 | Codex 0.162.0 | \n|---|---|---|\n| Root `oneOf` next to`properties` | works | property names lost | \n| 2,049 tools on one server | works | tools past 2,048 dropped | \n| One tool with no root `type` | whole server dropped | works | \n| int64 argument `1234567890123456789` | rounded | string or rounded | \n| Enum inside `anyOf` , schema over 5 KB | works | enum erased | \n| `prefixItems` tuple | works | declared as Array<string> | \n| `fetch.page` and`fetch_page` on one server | one misrouted | works | \n| `outputSchema` declared, text-only result | success turned into an error | works | \n\nThe probes also cleared some old suspects. Claude Code's root `anyOf` and 65-character name issues no longer reproduce at current versions, and neither does Codex's `resource_link` handling, so server authors can drop the workarounds for those.\n\n## How M3 runs each cell\n\nEvery cell is an ordinary pytest test that M3 runs against the real agent binary, at an exact version, while recording what crosses the wire. Here is the whole test for the `ids[]` case:\n\n``` python\n@pytest.mark.m3(suite_name=\"quirks\")\ndef test_q05_args_property_name_brackets(agent) -> None:\n    state = run_eval(agent, Path(__file__).resolve().parent)\n    assert state == \"works\", state\n```\n\n`run_eval` calls `agent.run(prompt, servers=(control, quirk))` and then checks the result with M3's `expect(result).to_have_tool_call(...)` against wire evidence. The `agent` fixture comes from the command line:\n\n```\nm3 test --runtime=managed \\\n  --harness [email protected]=claude-sonnet-5-5 \\\n  --harness [email protected]=gpt-5.6-sol \\\n  --trials 3 -- quirks/\n```\n\nA few things keep the cells honest:\n\n1. M3 downloads each agent release into a cache and records its version and SHA-256, so a new release gets its own column.\n2. M3 records separately whether the agent reported a tool call and whether the server received it. Most of the failures above happen between those two points.\n3. Each run puts a plain `echo` server next to the server under test. If`echo` fails, the cell is marked not measured, so a broken setup can't show up as an agent bug.\n4. The server generates a random marker for every trial, so the model can't get the answer from the prompt.\n5. We re-run every failure against the same binary without M3, to check that M3 itself isn't the cause.\n6. The table is rendered only from stored files, and re-rendering it gives the same bytes.\n\nThe full run took 66 executions and about 16 minutes. Claude Code averaged 9.8 seconds per execution; Codex averaged 18.2 seconds.\n\nAll twenty tests are in the M3 repository under [benchmarks/harness-quirks](https://github.com/sineframe/m3/tree/main/benchmarks/harness-quirks), with the command to run them yourself.\n\n## What surprised us\n\nOur first version found nothing. We started with five features taken from public bug reports, and all 36 executions passed on both agents. Those bugs had either been fixed or only showed up on setups we weren't testing, which is why we went back and probed the binaries directly.\n\nOne sentence in a prompt broke every Codex run. Our first prompts ended with \"If a tool is not available, say which one and stop.\" Codex took that as a reason to give up, and every cell came back control failed. Without the control server, we would have published that as a Codex bug.\n\nWith `gpt-5.6-sol`, Codex runs in code mode, so the model never sees your JSON Schema. It gets a compacted TypeScript declaration generated from it, and several of the Codex failures in this post happen in that translation.\n\n## Test your own MCP server on every agent\n\nThe matrix is just M3 tests, and you can write the same test for your own server in about ten lines. Install the CLI and scaffold a project:\n\n```\nuv tool install sf-m3-cli\nm3 init      # a skipped starter test and an .env.example\nm3 setup\n```\n\nWrite the behaviour you care about as a normal pytest test:\n\n``` python\nimport pytest\nfrom m3 import expect\n\n@pytest.mark.m3\ndef test_shipping(agent, shipping_server):\n    result = agent.run(\"Get a local shipping quote\", server=shipping_server)\n    expect(result).to_have_tool_call(\"shipping_quote\", server=shipping_server.name,\n                                     status=\"success\")\n```\n\nThen run it against as many agents and versions as you like:\n\n```\nm3 test --runtime=managed \\\n  --harness [email protected]=claude-sonnet-5-5 \\\n  --harness [email protected]=gpt-5.6-sol \\\n  --trials 3 -- tests/test_shipping.py\n```\n\nM3 ships with Claude Code, Codex, OpenCode and Pi built in, and any agent that speaks ACP can be added with a small manifest. Runs are saved locally and can be compared with a saved baseline, so when a new agent release breaks your server, you see a failing test before your users see the bug. To make that a pull-request check, see [gating an MCP server in CI](https://m3.sineframe.com/mcp-ci).\n\nM3 is open source under Apache 2.0: [github.com/sineframe/m3](https://github.com/sineframe/m3). Start with the [documentation](https://m3.sineframe.com/docs/). If your server has a quirk we should add to the matrix, open an issue there.", "url": "https://wpnews.pro/news/claude-code-and-codex-break-on-different-mcp-features", "canonical_source": "https://m3.sineframe.com/blog/claude-code-vs-codex-mcp", "published_at": "2026-10-10 15:16:02+00:00", "updated_at": "2026-10-10 15:47:19.964680+00:00", "lang": "en", "topics": ["ai-agents", "agent-protocols", "developer-tools", "ai-tools"], "entities": ["Claude Code", "Codex", "M3", "Anthropic", "OpenAI", "Harness Quirks Matrix"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/claude-code-and-codex-break-on-different-mcp-features", "markdown": "https://wpnews.pro/news/claude-code-and-codex-break-on-different-mcp-features.md", "text": "https://wpnews.pro/news/claude-code-and-codex-break-on-different-mcp-features.txt", "jsonld": "https://wpnews.pro/news/claude-code-and-codex-break-on-different-mcp-features.jsonld"}}