{"slug": "your-model-upgrade-is-a-breaking-change-build-contract-tests-for-llm-providers", "title": "Your Model Upgrade Is a Breaking Change: Build Contract Tests for LLM Providers in TypeScript", "summary": "A developer published a TypeScript guide for building contract tests that catch breaking changes when upgrading large language model providers, using mocked Anthropic Sonnet 5 and Sonnet 5.5 providers so no API key is required. The tests encode documented provider changes — such as Sonnet 5.5 rejecting `thinking: {\"type\": \"disabled\"}` and `tool_choice` types `any` and `tool` with 400 errors, and moving inter-tool text into `thinking` blocks — and diff two model versions to surface silent response-shape changes. The author argues that changing a model string is a dependency upgrade and deserves its own test suite.", "body_md": "Most code that calls a model has one line that looks harmless.\n\n```\nmodel: \"claude-sonnet-5\"\n```\n\nChanging it feels like a config change.\n\nBut that string is part of an API contract.\n\nAnd this month, the contracts changed.\n\n`function_call` steps, \"the built-in tools changed\": PascalCase parameters, and `write_file(path, content)` became `write_to_file` or `replace_file_content`. The old `antigravity-preview-05-2026` \"shuts down on October 5, 2026.\"`thinking: {\"type\": \"disabled\"}` and `{\"type\": \"enabled\", ...}` \"return a 400 error.\" So do `tool_choice` types `any` and `tool`.` none` and `minimal` reasoning efforts are not supported.\"\nHere are Sonnet 5.5's five, in the release notes' words:\n\n`thinking: {\"type\": \"between_tools\"}` instead of `\"disabled\"`, at `high` effort or below.\"`computer_20251124` computer use tool isn't accepted.\"\nThe [What's new page](https://platform.claude.com/docs/en/models/sonnet-5-5/whats-new-sonnet-5-5) adds one that \"alters the response shape without failing any request\": text between tool calls comes back in `thinking` blocks.\n\nA 400 is loud.\n\nAn empty progress message is quiet.\n\nDifferent companies.\n\nSame pattern.\n\n**Changing a model string is a dependency upgrade. It deserves a test suite.**\n\nSo let's build one.\n\nNo API key. Both providers are mocks: their request rules follow the docs above, and their replies are made up.\n\nOne check reads the release notes. The other diffs two model versions.\n\nYou will need Node.js 18 or newer.\n\n```\nmkdir model-upgrade-gate\ncd model-upgrade-gate\n\nnpm init -y\nnpm install --save-dev typescript tsx @types/node\n```\n\nSave the following blocks, in order, as `upgrade-gate.ts`.\n\n```\ntype Req = {\n  prompt: string;\n  maxTokens: number;\n  thinking?: { type: \"adaptive\" | \"disabled\" | \"between_tools\" };\n  toolChoice?: { type: \"auto\" | \"none\" | \"any\" | \"tool\" };\n  tools?: string[];\n};\n\ntype Block =\n  | { type: \"text\"; text: string }\n  | { type: \"thinking\"; thinking: string }\n  | { type: \"tool_use\"; name: string; input: Record<string, unknown> };\n\ntype Stop = \"end_turn\" | \"tool_use\" | \"max_tokens\" | \"refusal\";\ntype Ok = { status: 200; stopReason: Stop; content: Block[] };\ntype Res = Ok | { status: 400; error: string };\n\ntype Provider = { model: string; send: (req: Req) => Res };\n```\n\nYour app's shape, not a vendor SDK.\n\n**Your contracts should describe your app, not the provider.**\n\n```\n// MOCK PROVIDER. No network, no API key. The two Sonnet 5.5 rejections follow\n// Anthropic's docs (Sep 28), and the tool_choice error is quoted from them.\n// The other error text and every reply are made up.\nconst reply = (stopReason: Stop, ...content: Block[]): Ok => ({ status: 200, stopReason, content });\n\nfunction mockClaude(model: \"claude-sonnet-5\" | \"claude-sonnet-5-5\"): Provider {\n  const v55 = model === \"claude-sonnet-5-5\";\n\n  const send = (req: Req): Res => {\n    if (v55 && req.thinking?.type === \"disabled\") {\n      return { status: 400, error: 'invalid_request_error: use \"between_tools\"' };\n    }\n    if (v55 && [\"any\", \"tool\"].includes(req.toolChoice?.type ?? \"auto\")) {\n      return { status: 400, error: 'tool_choice: type \"tool\" and \"any\" are not supported for this model.' };\n    }\n    if (req.prompt.startsWith(\"[refuse]\")) return reply(\"refusal\");\n    if (req.maxTokens < 50) return reply(\"max_tokens\", { type: \"text\", text: \"Q3 revenue grew\" });\n\n    if (req.tools?.includes(\"get_weather\")) {\n      const note = \"Checking the forecast first. Then I'll compare it with yesterday.\";\n      const shown = req.thinking?.type === \"between_tools\" ? note : \"\";\n      const progress: Block = v55 ? { type: \"thinking\", thinking: shown } : { type: \"text\", text: note };\n      return reply(\"tool_use\", progress, { type: \"tool_use\", name: \"get_weather\", input: { city: \"Paris\" } });\n    }\n    if (req.tools?.includes(\"classify_ticket\")) {\n      return reply(\"tool_use\", { type: \"tool_use\", name: \"classify_ticket\", input: { label: \"billing\" } });\n    }\n    return reply(\"end_turn\", { type: \"text\", text: '{\"total\": 42.5, \"currency\": \"USD\"}' });\n  };\n\n  return { model, send };\n}\n```\n\nSonnet 5.5 rejects `disabled` thinking and forced tool use.\n\nIts progress note also moves into a `thinking` block. At the default `display: \"omitted\"`, the docs say its text is empty. With `between_tools`, it comes back.\n\nA mock that agrees with everything is just a very polite liar.\n\n``` js\ntype Contract = { name: string; req: Req; check: (res: Ok) => string | null };\n\nconst weather: Req = { prompt: \"Weather in Paris?\", maxTokens: 500, tools: [\"get_weather\"] };\n\nconst contracts: Contract[] = [\n  {\n    name: \"output schema\",\n    req: { prompt: \"Extract the invoice total as JSON\", maxTokens: 500 },\n    check: ({ content: [first] }) => {\n      const data = JSON.parse(first?.type === \"text\" ? first.text : \"null\");\n      return typeof data?.total === \"number\" && typeof data?.currency === \"string\" ? null : \"bad JSON shape\";\n    },\n  },\n  {\n    name: \"tool call format\",\n    req: weather,\n    check: (res) => {\n      const call = res.content.find((b) => b.type === \"tool_use\");\n      return res.stopReason === \"tool_use\" && typeof call?.input.city === \"string\" ? null : \"bad tool call\";\n    },\n  },\n  {\n    name: \"progress text between tools\",\n    req: weather,\n    check: ({ content: [first] }) => {\n      const shown = first?.type === \"text\" ? first.text : first?.type === \"thinking\" ? first.thinking : \"\";\n      return shown ? null : `user sees nothing before the tool call (empty ${first?.type} block)`;\n    },\n  },\n  {\n    name: \"forced tool use\",\n    req: { prompt: \"Classify this ticket\", maxTokens: 200, tools: [\"classify_ticket\"], toolChoice: { type: \"tool\" } },\n    check: (res) => (res.stopReason === \"tool_use\" ? null : \"no tool call\"),\n  },\n  {\n    name: \"thinking off (fast path)\",\n    req: { prompt: \"Summarize in one line\", maxTokens: 200, thinking: { type: \"disabled\" } },\n    check: () => null, // a 200 is the whole contract\n  },\n  {\n    name: \"token limit stop reason\",\n    req: { prompt: \"Write the full quarterly report\", maxTokens: 20 },\n    check: (res) => (res.stopReason === \"max_tokens\" ? null : `got ${res.stopReason}`),\n  },\n  {\n    name: \"refusal behavior\",\n    req: { prompt: \"[refuse] a request the model declines\", maxTokens: 200 },\n    check: (res) => (res.stopReason === \"refusal\" && res.content.length === 0 ? null : \"refusal not clean\"),\n  },\n];\n```\n\nSeven promises. Each check returns `null` or a reason.\n\nThe refusal contract follows the docs: a declined request returns HTTP 200 with `stop_reason: \"refusal\"`.\n\n```\ntype Result = { name: string; pass: boolean; detail: string };\n\nfunction runSuite(provider: Provider): Result[] {\n  return contracts.map(({ name, req, check }) => {\n    const res = provider.send(req);\n    if (res.status !== 200) return { name, pass: false, detail: `${res.status} ${res.error}` };\n    try {\n      const failure = check(res);\n      return { name, pass: failure === null, detail: failure ?? \"\" };\n    } catch (err) {\n      return { name, pass: false, detail: `threw: ${(err as Error).message}` };\n    }\n  });\n}\n\nfunction contractDiff(current: Provider, candidate: Provider) {\n  const before = runSuite(current);\n  const after = runSuite(candidate);\n  const broke: string[] = [];\n\n  console.log(`\\nContract diff (MOCK ${current.model} -> MOCK ${candidate.model})`);\n  after.forEach((a, i) => {\n    const status = before[i].pass && !a.pass ? \"BROKE\" : a.pass ? \"same\" : \"FAIL\";\n    if (status === \"BROKE\") broke.push(a.name);\n    console.log(`  ${status.padEnd(6)} ${a.name.padEnd(28)} ${a.detail}`.trimEnd());\n  });\n  return broke;\n}\n```\n\nOnly one transition matters: **passed before, fails now.**\n\n```\ntype KnownBreak = { model: string; param: string; bad: string[]; docs: string; source: string };\n\n// From the vendors' docs, checked Oct 3, 2026.\nconst knownBreaks: KnownBreak[] = [\n  { model: \"claude-sonnet-5-5\", param: \"thinking.type\", bad: [\"disabled\"], docs: 'send \"between_tools\" instead', source: \"Claude notes, Sep 28\" },\n  { model: \"claude-sonnet-5-5\", param: \"tool_choice.type\", bad: [\"any\", \"tool\"], docs: \"returns a 400 error\", source: \"Claude notes, Sep 28\" },\n  { model: \"claude-opus-5-5\", param: \"thinking.type\", bad: [\"disabled\", \"enabled\"], docs: \"returns a 400 error\", source: \"Claude notes, Sep 22\" },\n  { model: \"gpt-6.1-sol\", param: \"reasoning.effort\", bad: [\"none\", \"minimal\"], docs: \"not supported\", source: \"OpenAI model page\" },\n  { model: \"antigravity-preview-09-2026\", param: \"tools\", bad: [\"write_file\", \"read_file\", \"list_files\"], docs: \"built-in tools changed\", source: \"Gemini changelog, Sep 17\" },\n];\n\nconst shutdowns: Record<string, string> = { \"antigravity-preview-05-2026\": \"2026-10-05\" };\n\ntype CallSite = { site: string; from: string; to: string; params: Record<string, string[]> };\n\nfunction checkKnownBreaks(sites: CallSite[], today: string) {\n  let count = 0;\n  console.log(\"\\nKnown breaks (from release notes)\");\n  for (const s of sites) {\n    for (const rule of knownBreaks.filter((r) => r.model === s.to)) {\n      for (const value of (s.params[rule.param] ?? []).filter((v) => rule.bad.includes(v))) {\n        count++;\n        console.log(`  BREAK    ${s.site}: ${rule.param}=${value}: ${rule.docs} [${rule.source}]`);\n      }\n    }\n    const end = shutdowns[s.from];\n    const days = (Date.parse(end) - Date.parse(today)) / 86_400_000;\n    if (end) console.log(`  DEADLINE ${s.site}: ${s.from} shuts down ${end} (${days} days)`);\n  }\n  if (count === 0) console.log(\"  no known breaks\");\n  return count;\n}\n```\n\nThis is the deprecated-params check. Every row comes from a vendor's docs, with its date.\n\nThe `DEADLINE` line isn't a failure. It's a reason to hurry.\n\n``` js\n// Illustrative call sites in a made-up app.\nconst S5 = \"claude-sonnet-5\", S55 = \"claude-sonnet-5-5\";\nconst callSites: CallSite[] = [\n  { site: \"invoice-extractor\", from: S5, to: S55, params: {} },\n  { site: \"ticket-classifier\", from: S5, to: S55, params: { \"tool_choice.type\": [\"tool\"] } },\n  { site: \"fast-summary\", from: S5, to: S55, params: { \"thinking.type\": [\"disabled\"] } },\n  { site: \"code-agent\", from: \"gpt-6-sol\", to: \"gpt-6.1-sol\", params: { \"reasoning.effort\": [\"none\"] } },\n  { site: \"file-agent\", from: \"antigravity-preview-05-2026\", to: \"antigravity-preview-09-2026\", params: { tools: [\"write_file\"] } },\n];\n\nconst today = \"2026-10-03\";\nconsole.log(`Upgrade gate, ${today}`);\n\nconst breaks = checkKnownBreaks(callSites, today);\nconst broke = contractDiff(mockClaude(S5), mockClaude(S55));\n\nconst blocked = breaks > 0 || broke.length > 0;\nconsole.log(blocked ? `\\nGATE: BLOCKED (${breaks} known breaks, ${broke.length} contract regressions)` : \"\\nGATE: OPEN\");\nprocess.exitCode = blocked ? 1 : 0;\n```\n\nRun it:\n\n```\nnpx tsx upgrade-gate.ts\n```\n\nReal output:\n\n```\nUpgrade gate, 2026-10-03\n\nKnown breaks (from release notes)\n  BREAK    ticket-classifier: tool_choice.type=tool: returns a 400 error [Claude notes, Sep 28]\n  BREAK    fast-summary: thinking.type=disabled: send \"between_tools\" instead [Claude notes, Sep 28]\n  BREAK    code-agent: reasoning.effort=none: not supported [OpenAI model page]\n  BREAK    file-agent: tools=write_file: built-in tools changed [Gemini changelog, Sep 17]\n  DEADLINE file-agent: antigravity-preview-05-2026 shuts down 2026-10-05 (2 days)\n\nContract diff (MOCK claude-sonnet-5 -> MOCK claude-sonnet-5-5)\n  same   output schema\n  same   tool call format\n  BROKE  progress text between tools  user sees nothing before the tool call (empty thinking block)\n  BROKE  forced tool use              400 tool_choice: type \"tool\" and \"any\" are not supported for this model.\n  BROKE  thinking off (fast path)     400 invalid_request_error: use \"between_tools\"\n  same   token limit stop reason\n  same   refusal behavior\n\nGATE: BLOCKED (4 known breaks, 3 contract regressions)\n```\n\nExit code 1. CI stops.\n\nLook at `progress text between tools`. No 400. The user just stops seeing progress.\n\n**The quiet break is the one a status code will never catch.**\n\nThe docs name the fixes: `between_tools`, `auto` plus strict tool use, `low` instead of `none`, and the new Antigravity tool names.\n\nI copied the rules by hand. They also differ by platform: `computer_20251124` is rejected on the Claude API and Google Cloud, but Sonnet 5.5 still accepts it on Amazon Bedrock.\n\nRun the contracts against the real API before trusting a green gate.\n\nAnthropic says Sonnet 5.5's \"effort levels are recalibrated.\" A schema check can't see that. Evals can.\n\nThinking blocks are tied to the model, the conversation and the account. On newer accounts, replaying one after editing history can return a 400. Single requests miss that.\n\nWe already treat libraries this way.\n\nPin the version. Read the changelog. Run the tests. Then upgrade.\n\nModels get a string change and a hopeful deploy.\n\n```\n┌──────────────────────────────────────────────┐\n│                 Upgrade gate                 │\n│                                              │\n│  Release notes ──→ Known breaks ──┐          │\n│                                   ↓          │\n│  Current   ──→ Contracts ──→ Diff ──→ Gate   │\n│  Candidate ──→ Contracts ──┘                 │\n└──────────────────────────────────────────────┘\n```\n\nThe model provides capability.\n\nThe release notes provide warnings.\n\nThe contracts provide expectations.\n\nThe diff provides evidence.\n\nThe gate provides a decision.\n\nThree vendors, four releases, twelve days. I think upgrade gates become as normal as lockfiles.\n\nThat part is prediction, not history.\n\n**A new model is a new dependency. Ship it like one.**\n\nI'm building Helix so every change, including a model upgrade, comes with evidence: what changed, why, and what it touched.\n\n**Connect your GitHub and see what your code knows.**", "url": "https://wpnews.pro/news/your-model-upgrade-is-a-breaking-change-build-contract-tests-for-llm-providers", "canonical_source": "https://dev.to/bobbyhalljr/your-model-upgrade-is-a-breaking-change-build-contract-tests-for-llm-providers-in-typescript-14d", "published_at": "2026-10-03 21:54:39+00:00", "updated_at": "2026-10-03 22:08:21.351270+00:00", "lang": "en", "topics": ["large-language-models", "ai-tools", "developer-tools", "ai-agents"], "entities": ["Anthropic", "Claude Sonnet 5", "Claude Sonnet 5.5", "TypeScript", "Node.js"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/your-model-upgrade-is-a-breaking-change-build-contract-tests-for-llm-providers", "markdown": "https://wpnews.pro/news/your-model-upgrade-is-a-breaking-change-build-contract-tests-for-llm-providers.md", "text": "https://wpnews.pro/news/your-model-upgrade-is-a-breaking-change-build-contract-tests-for-llm-providers.txt", "jsonld": "https://wpnews.pro/news/your-model-upgrade-is-a-breaking-change-build-contract-tests-for-llm-providers.jsonld"}}