# Your Model Upgrade Is a Breaking Change: Build Contract Tests for LLM Providers in TypeScript

> Source: <https://dev.to/bobbyhalljr/your-model-upgrade-is-a-breaking-change-build-contract-tests-for-llm-providers-in-typescript-14d>
> Published: 2026-10-03 21:54:39+00:00

Most code that calls a model has one line that looks harmless.

```
model: "claude-sonnet-5"
```

Changing it feels like a config change.

But that string is part of an API contract.

And this month, the contracts changed.

`function_call` steps, "the built-in tools changed": PascalCase parameters, and `write_file(path, content)` became `write_to_file` or `replace_file_content`. The old `antigravity-preview-05-2026` "shuts down on October 5, 2026."`thinking: {"type": "disabled"}` and `{"type": "enabled", ...}` "return a 400 error." So do `tool_choice` types `any` and `tool`.` none` and `minimal` reasoning efforts are not supported."
Here are Sonnet 5.5's five, in the release notes' words:

`thinking: {"type": "between_tools"}` instead of `"disabled"`, at `high` effort or below."`computer_20251124` computer use tool isn't accepted."
The [What's new page](https://platform.claude.com/docs/en/models/sonnet-5-5/whats-new-sonnet-5-5) adds one that "alters the response shape without failing any request": text between tool calls comes back in `thinking` blocks.

A 400 is loud.

An empty progress message is quiet.

Different companies.

Same pattern.

**Changing a model string is a dependency upgrade. It deserves a test suite.**

So let's build one.

No API key. Both providers are mocks: their request rules follow the docs above, and their replies are made up.

One check reads the release notes. The other diffs two model versions.

You will need Node.js 18 or newer.

```
mkdir model-upgrade-gate
cd model-upgrade-gate

npm init -y
npm install --save-dev typescript tsx @types/node
```

Save the following blocks, in order, as `upgrade-gate.ts`.

```
type Req = {
  prompt: string;
  maxTokens: number;
  thinking?: { type: "adaptive" | "disabled" | "between_tools" };
  toolChoice?: { type: "auto" | "none" | "any" | "tool" };
  tools?: string[];
};

type Block =
  | { type: "text"; text: string }
  | { type: "thinking"; thinking: string }
  | { type: "tool_use"; name: string; input: Record<string, unknown> };

type Stop = "end_turn" | "tool_use" | "max_tokens" | "refusal";
type Ok = { status: 200; stopReason: Stop; content: Block[] };
type Res = Ok | { status: 400; error: string };

type Provider = { model: string; send: (req: Req) => Res };
```

Your app's shape, not a vendor SDK.

**Your contracts should describe your app, not the provider.**

```
// MOCK PROVIDER. No network, no API key. The two Sonnet 5.5 rejections follow
// Anthropic's docs (Sep 28), and the tool_choice error is quoted from them.
// The other error text and every reply are made up.
const reply = (stopReason: Stop, ...content: Block[]): Ok => ({ status: 200, stopReason, content });

function mockClaude(model: "claude-sonnet-5" | "claude-sonnet-5-5"): Provider {
  const v55 = model === "claude-sonnet-5-5";

  const send = (req: Req): Res => {
    if (v55 && req.thinking?.type === "disabled") {
      return { status: 400, error: 'invalid_request_error: use "between_tools"' };
    }
    if (v55 && ["any", "tool"].includes(req.toolChoice?.type ?? "auto")) {
      return { status: 400, error: 'tool_choice: type "tool" and "any" are not supported for this model.' };
    }
    if (req.prompt.startsWith("[refuse]")) return reply("refusal");
    if (req.maxTokens < 50) return reply("max_tokens", { type: "text", text: "Q3 revenue grew" });

    if (req.tools?.includes("get_weather")) {
      const note = "Checking the forecast first. Then I'll compare it with yesterday.";
      const shown = req.thinking?.type === "between_tools" ? note : "";
      const progress: Block = v55 ? { type: "thinking", thinking: shown } : { type: "text", text: note };
      return reply("tool_use", progress, { type: "tool_use", name: "get_weather", input: { city: "Paris" } });
    }
    if (req.tools?.includes("classify_ticket")) {
      return reply("tool_use", { type: "tool_use", name: "classify_ticket", input: { label: "billing" } });
    }
    return reply("end_turn", { type: "text", text: '{"total": 42.5, "currency": "USD"}' });
  };

  return { model, send };
}
```

Sonnet 5.5 rejects `disabled` thinking and forced tool use.

Its progress note also moves into a `thinking` block. At the default `display: "omitted"`, the docs say its text is empty. With `between_tools`, it comes back.

A mock that agrees with everything is just a very polite liar.

``` js
type Contract = { name: string; req: Req; check: (res: Ok) => string | null };

const weather: Req = { prompt: "Weather in Paris?", maxTokens: 500, tools: ["get_weather"] };

const contracts: Contract[] = [
  {
    name: "output schema",
    req: { prompt: "Extract the invoice total as JSON", maxTokens: 500 },
    check: ({ content: [first] }) => {
      const data = JSON.parse(first?.type === "text" ? first.text : "null");
      return typeof data?.total === "number" && typeof data?.currency === "string" ? null : "bad JSON shape";
    },
  },
  {
    name: "tool call format",
    req: weather,
    check: (res) => {
      const call = res.content.find((b) => b.type === "tool_use");
      return res.stopReason === "tool_use" && typeof call?.input.city === "string" ? null : "bad tool call";
    },
  },
  {
    name: "progress text between tools",
    req: weather,
    check: ({ content: [first] }) => {
      const shown = first?.type === "text" ? first.text : first?.type === "thinking" ? first.thinking : "";
      return shown ? null : `user sees nothing before the tool call (empty ${first?.type} block)`;
    },
  },
  {
    name: "forced tool use",
    req: { prompt: "Classify this ticket", maxTokens: 200, tools: ["classify_ticket"], toolChoice: { type: "tool" } },
    check: (res) => (res.stopReason === "tool_use" ? null : "no tool call"),
  },
  {
    name: "thinking off (fast path)",
    req: { prompt: "Summarize in one line", maxTokens: 200, thinking: { type: "disabled" } },
    check: () => null, // a 200 is the whole contract
  },
  {
    name: "token limit stop reason",
    req: { prompt: "Write the full quarterly report", maxTokens: 20 },
    check: (res) => (res.stopReason === "max_tokens" ? null : `got ${res.stopReason}`),
  },
  {
    name: "refusal behavior",
    req: { prompt: "[refuse] a request the model declines", maxTokens: 200 },
    check: (res) => (res.stopReason === "refusal" && res.content.length === 0 ? null : "refusal not clean"),
  },
];
```

Seven promises. Each check returns `null` or a reason.

The refusal contract follows the docs: a declined request returns HTTP 200 with `stop_reason: "refusal"`.

```
type Result = { name: string; pass: boolean; detail: string };

function runSuite(provider: Provider): Result[] {
  return contracts.map(({ name, req, check }) => {
    const res = provider.send(req);
    if (res.status !== 200) return { name, pass: false, detail: `${res.status} ${res.error}` };
    try {
      const failure = check(res);
      return { name, pass: failure === null, detail: failure ?? "" };
    } catch (err) {
      return { name, pass: false, detail: `threw: ${(err as Error).message}` };
    }
  });
}

function contractDiff(current: Provider, candidate: Provider) {
  const before = runSuite(current);
  const after = runSuite(candidate);
  const broke: string[] = [];

  console.log(`\nContract diff (MOCK ${current.model} -> MOCK ${candidate.model})`);
  after.forEach((a, i) => {
    const status = before[i].pass && !a.pass ? "BROKE" : a.pass ? "same" : "FAIL";
    if (status === "BROKE") broke.push(a.name);
    console.log(`  ${status.padEnd(6)} ${a.name.padEnd(28)} ${a.detail}`.trimEnd());
  });
  return broke;
}
```

Only one transition matters: **passed before, fails now.**

```
type KnownBreak = { model: string; param: string; bad: string[]; docs: string; source: string };

// From the vendors' docs, checked Oct 3, 2026.
const knownBreaks: KnownBreak[] = [
  { model: "claude-sonnet-5-5", param: "thinking.type", bad: ["disabled"], docs: 'send "between_tools" instead', source: "Claude notes, Sep 28" },
  { model: "claude-sonnet-5-5", param: "tool_choice.type", bad: ["any", "tool"], docs: "returns a 400 error", source: "Claude notes, Sep 28" },
  { model: "claude-opus-5-5", param: "thinking.type", bad: ["disabled", "enabled"], docs: "returns a 400 error", source: "Claude notes, Sep 22" },
  { model: "gpt-6.1-sol", param: "reasoning.effort", bad: ["none", "minimal"], docs: "not supported", source: "OpenAI model page" },
  { model: "antigravity-preview-09-2026", param: "tools", bad: ["write_file", "read_file", "list_files"], docs: "built-in tools changed", source: "Gemini changelog, Sep 17" },
];

const shutdowns: Record<string, string> = { "antigravity-preview-05-2026": "2026-10-05" };

type CallSite = { site: string; from: string; to: string; params: Record<string, string[]> };

function checkKnownBreaks(sites: CallSite[], today: string) {
  let count = 0;
  console.log("\nKnown breaks (from release notes)");
  for (const s of sites) {
    for (const rule of knownBreaks.filter((r) => r.model === s.to)) {
      for (const value of (s.params[rule.param] ?? []).filter((v) => rule.bad.includes(v))) {
        count++;
        console.log(`  BREAK    ${s.site}: ${rule.param}=${value}: ${rule.docs} [${rule.source}]`);
      }
    }
    const end = shutdowns[s.from];
    const days = (Date.parse(end) - Date.parse(today)) / 86_400_000;
    if (end) console.log(`  DEADLINE ${s.site}: ${s.from} shuts down ${end} (${days} days)`);
  }
  if (count === 0) console.log("  no known breaks");
  return count;
}
```

This is the deprecated-params check. Every row comes from a vendor's docs, with its date.

The `DEADLINE` line isn't a failure. It's a reason to hurry.

``` js
// Illustrative call sites in a made-up app.
const S5 = "claude-sonnet-5", S55 = "claude-sonnet-5-5";
const callSites: CallSite[] = [
  { site: "invoice-extractor", from: S5, to: S55, params: {} },
  { site: "ticket-classifier", from: S5, to: S55, params: { "tool_choice.type": ["tool"] } },
  { site: "fast-summary", from: S5, to: S55, params: { "thinking.type": ["disabled"] } },
  { site: "code-agent", from: "gpt-6-sol", to: "gpt-6.1-sol", params: { "reasoning.effort": ["none"] } },
  { site: "file-agent", from: "antigravity-preview-05-2026", to: "antigravity-preview-09-2026", params: { tools: ["write_file"] } },
];

const today = "2026-10-03";
console.log(`Upgrade gate, ${today}`);

const breaks = checkKnownBreaks(callSites, today);
const broke = contractDiff(mockClaude(S5), mockClaude(S55));

const blocked = breaks > 0 || broke.length > 0;
console.log(blocked ? `\nGATE: BLOCKED (${breaks} known breaks, ${broke.length} contract regressions)` : "\nGATE: OPEN");
process.exitCode = blocked ? 1 : 0;
```

Run it:

```
npx tsx upgrade-gate.ts
```

Real output:

```
Upgrade gate, 2026-10-03

Known breaks (from release notes)
  BREAK    ticket-classifier: tool_choice.type=tool: returns a 400 error [Claude notes, Sep 28]
  BREAK    fast-summary: thinking.type=disabled: send "between_tools" instead [Claude notes, Sep 28]
  BREAK    code-agent: reasoning.effort=none: not supported [OpenAI model page]
  BREAK    file-agent: tools=write_file: built-in tools changed [Gemini changelog, Sep 17]
  DEADLINE file-agent: antigravity-preview-05-2026 shuts down 2026-10-05 (2 days)

Contract diff (MOCK claude-sonnet-5 -> MOCK claude-sonnet-5-5)
  same   output schema
  same   tool call format
  BROKE  progress text between tools  user sees nothing before the tool call (empty thinking block)
  BROKE  forced tool use              400 tool_choice: type "tool" and "any" are not supported for this model.
  BROKE  thinking off (fast path)     400 invalid_request_error: use "between_tools"
  same   token limit stop reason
  same   refusal behavior

GATE: BLOCKED (4 known breaks, 3 contract regressions)
```

Exit code 1. CI stops.

Look at `progress text between tools`. No 400. The user just stops seeing progress.

**The quiet break is the one a status code will never catch.**

The docs name the fixes: `between_tools`, `auto` plus strict tool use, `low` instead of `none`, and the new Antigravity tool names.

I copied the rules by hand. They also differ by platform: `computer_20251124` is rejected on the Claude API and Google Cloud, but Sonnet 5.5 still accepts it on Amazon Bedrock.

Run the contracts against the real API before trusting a green gate.

Anthropic says Sonnet 5.5's "effort levels are recalibrated." A schema check can't see that. Evals can.

Thinking blocks are tied to the model, the conversation and the account. On newer accounts, replaying one after editing history can return a 400. Single requests miss that.

We already treat libraries this way.

Pin the version. Read the changelog. Run the tests. Then upgrade.

Models get a string change and a hopeful deploy.

```
┌──────────────────────────────────────────────┐
│                 Upgrade gate                 │
│                                              │
│  Release notes ──→ Known breaks ──┐          │
│                                   ↓          │
│  Current   ──→ Contracts ──→ Diff ──→ Gate   │
│  Candidate ──→ Contracts ──┘                 │
└──────────────────────────────────────────────┘
```

The model provides capability.

The release notes provide warnings.

The contracts provide expectations.

The diff provides evidence.

The gate provides a decision.

Three vendors, four releases, twelve days. I think upgrade gates become as normal as lockfiles.

That part is prediction, not history.

**A new model is a new dependency. Ship it like one.**

I'm building Helix so every change, including a model upgrade, comes with evidence: what changed, why, and what it touched.

**Connect your GitHub and see what your code knows.**
