GLM 5.3 signed my commit as Claude Fable 5. The harness is in the weights A September 25, 2026 analysis of Stencil's hashline edit-tool research found that the +15-point coding gains attributed to the harness are mostly baked into model post-training, with hashline versus Claude-style str_replace averaging only about +3.4 points across 16 models, from +11.3 (Claude Haiku 4.5, Gemini 3 Flash) to -8.3 (DeepSeek V3.2). Stencil's original February experiment replaced the edit tool with hashline, in which every line a model reads is tagged with a short content hash, and reported Grok Code Fast 1 rising from 6.7% to 68.3% pass rates. The author argues the best edit tool for a model is largely decided by whatever harness the lab used for reinforcement learning, so third-party harnesses operate as a point in the training distribution rather than a neutral contract. The harness is in the weights September 25, 2026 In February, Stencil published We improved 15 LLMs at coding in one afternoon. Only the harness changed https://stencil.so/blog/the-harness-problem . They replaced the edit tool with "hashline": every line a model reads comes back tagged with a short content hash, and edits point at those tags instead of repeating old text. Pass rates went up across 16 models. Grok Code Fast 1 went from 6.7% to 68.3%. The post ended with: "The model is the moat. The harness is the bridge." After seven months of replications, I think that line is half wrong. The bridge is part of the moat. The best edit tool for a model is mostly decided in post-training, by whatever harness the lab used for RL, and the labs have every reason to keep it that way. Tool calls are text The model gets a transcript, a system prompt and a list of tool definitions. The server flattens all of it into one long prompt with special marker tokens. At some point the model emits a span that the API reads as "call this tool with these arguments." It emits that span because it was trained and rewarded on examples of that exact format. Armin Ronacher's Better Models: Worse Tools https://lucumr.pocoo.org/2026/7/4/better-models-worse-tools/ shows what this looks like. Anthropic's serialization isn't public, but the markers that have leaked look like pseudo-XML: