AI coding agents still write your SDK's old API — so I built a type-checker to measure it A developer built SDKProof, a tool that uses TypeScript's compiler (tsc --noEmit) to measure how often AI coding agents write code against outdated library APIs. Testing Claude Opus 4 on three SDKs, it found scores of 80/100 for Prisma 7, 90/100 for Vercel AI SDK 7, and 90/100 for Zod 4, with errors stemming from renamed options like 'parameters' to 'inputSchema'. Here's a bug I kept hitting. I'd ask an AI assistant to write some code against a library — Prisma, the Vercel AI SDK, Zod — and the code would look completely right. Clean, idiomatic, exactly the shape I expected. Then I'd run it, and it wouldn't compile. The reason was always the same: the library had shipped a new major version, and the model was writing the previous major's API from memory. parameters instead of inputSchema . required error instead of error . A new PrismaClient { datasources } call that no longer exists. Small things — but enough to break the build on the first try. Models are frozen at their training cutoff. Libraries are not. So there's a gap between "the API the model reaches for" and "the API you actually have installed" — and that gap is widest right after a library ships a breaking change. I wanted to know: how big is that gap, exactly — and can I measure it objectively? So I built SDKProof https://sdkproof.dev . The trick to measuring this without hand-waving is to not use an LLM to grade an LLM. Instead: tsc --noEmit .No LLM judge. No "looks plausible." The installed package's type definitions are the ground truth, and tsc is the referee. A pass means the code would actually build against the version you have. One design detail that matters: the prompts name the functions but never the option names . I ask for "a tool with a description and an input schema," not "use inputSchema ." That way I'm measuring what the model naturally reaches for — not whether it can echo back a name I already handed it. Running claude-opus-4-8 across three SDKs: | SDK | Package | Score | Where it breaks | |---|---|---|---| | Prisma 7 | @prisma/client | 80 / 100 | still writes removed v6 setup — new PrismaClient with datasources , $use middleware | | Vercel AI SDK 7 | ai | 90 / 100 | old tool wiring — parameters now inputSchema , removed maxSteps | | Zod 4 | zod | 90 / 100 | removed required error now error — but nails the new 2-arg z.record | Here's a representative miss. Ask for a tool definition with the AI SDK and you tend to get this: js import { tool } from 'ai'; import { z } from 'zod'; const weatherTool = tool { description: 'Get the weather for a city', parameters: z.object { city: z.string } , // ✗ this was the v4 API execute: async { city } = getWeather city , } ; Looks right. But against ai v7, tsc says: error TS2353: Object literal may only specify known properties, and 'parameters' does not exist in type 'Tool<... '. Because the option was renamed to inputSchema : js const weatherTool = tool { description: 'Get the weather for a city', inputSchema: z.object { city: z.string } , // ✓ v5+ execute: async { city } = getWeather city , } ; Everything else about the code is fine. It's one renamed key — and it's exactly the kind of thing that slips past a quick read but stops the build cold. The interesting part isn't any single score — it's why they differ. Notice the newest breaking change scores worst. The Vercel AI SDK and Zod shipped their renames in 2025; by now the model has largely absorbed them and mostly gets them right 90/100 . Prisma 7 is more recent, and the model hasn't caught up 80/100 — it still writes setup calls that were removed. That's the whole thesis: a model's "readiness" for an SDK tracks how recently that SDK changed. Which means this isn't a one-time audit. The gap: So it's something to monitor , not measure once. A library that scores 95 today can drop to 70 the week it ships v-next — then climb back as the next model generation learns it. I'd rather you trust the number than oversell it: None of that breaks the signal — it just scopes it. "Does the model reach for API surface that still exists?" is a real, useful question, and the compiler answers it without opinion. Adding an SDK is about an afternoon: install the package, add a tsconfig , write a tasks file, register it. The harness handles generation, type-checking, and scoring. I'm looking for the next SDKs to score. What library have you watched an AI assistant get wrong since its last major? Tell me and I'll run it.