cd /news/artificial-intelligence/ai-coding-agents-still-write-your-sd… · home topics artificial-intelligence article
[ARTICLE · art-66048] src=dev.to ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

AI coding agents still write your SDK's old API — so I built a type-checker to measure it

A developer built SDKProof, a tool that uses TypeScript's compiler (tsc --noEmit) to measure how often AI coding agents write code against outdated library APIs. Testing Claude Opus 4 on three SDKs, it found scores of 80/100 for Prisma 7, 90/100 for Vercel AI SDK 7, and 90/100 for Zod 4, with errors stemming from renamed options like 'parameters' to 'inputSchema'.

read4 min views6 publishedJul 20, 2026

Here's a bug I kept hitting. I'd ask an AI assistant to write some code against a library — Prisma, the Vercel AI SDK, Zod — and the code would look completely right. Clean, idiomatic, exactly the shape I expected. Then I'd run it, and it wouldn't compile.

The reason was always the same: the library had shipped a new major version, and the model was writing the previous major's API from memory. parameters

instead of inputSchema

. required_error

instead of error

. A new PrismaClient({ datasources })

call that no longer exists. Small things — but enough to break the build on the first try.

Models are frozen at their training cutoff. Libraries are not. So there's a gap between "the API the model reaches for" and "the API you actually have installed" — and that gap is widest right after a library ships a breaking change.

I wanted to know: how big is that gap, exactly — and can I measure it objectively? So I built SDKProof.

The trick to measuring this without hand-waving is to not use an LLM to grade an LLM. Instead:

tsc --noEmit

.No LLM judge. No "looks plausible." The installed package's type definitions are the ground truth, and tsc

is the referee. A pass means the code would actually build against the version you have.

One design detail that matters: the prompts name the functions but never the option names. I ask for "a tool with a description and an input schema," not "use inputSchema

." That way I'm measuring what the model naturally reaches for — not whether it can echo back a name I already handed it.

Running claude-opus-4-8

across three SDKs:

SDK Package Score Where it breaks
Prisma 7 @prisma/client
80 / 100
still writes removed v6 setup — new PrismaClient() with datasources , $use middleware
Vercel AI SDK 7 ai
90 / 100
old tool wiring — parameters (now inputSchema ), removed maxSteps
Zod 4 zod
90 / 100
removed required_error (now error ) — but nails the new 2-arg z.record()

Here's a representative miss. Ask for a tool definition with the AI SDK and you tend to get this:

import { tool } from 'ai';
import { z } from 'zod';

const weatherTool = tool({
  description: 'Get the weather for a city',
  parameters: z.object({ city: z.string() }), // ✗ this was the v4 API
  execute: async ({ city }) => getWeather(city),
});

Looks right. But against ai

v7, tsc

says:

error TS2353: Object literal may only specify known properties,
and 'parameters' does not exist in type 'Tool<...>'.

Because the option was renamed to inputSchema

:

const weatherTool = tool({
  description: 'Get the weather for a city',
  inputSchema: z.object({ city: z.string() }), // ✓ v5+
  execute: async ({ city }) => getWeather(city),
});

Everything else about the code is fine. It's one renamed key — and it's exactly the kind of thing that slips past a quick read but stops the build cold.

The interesting part isn't any single score — it's why they differ.

Notice the newest breaking change scores worst. The Vercel AI SDK and Zod shipped their renames in 2025; by now the model has largely absorbed them and mostly gets them right (90/100). Prisma 7 is more recent, and the model hasn't caught up (80/100) — it still writes setup calls that were removed.

That's the whole thesis: a model's "readiness" for an SDK tracks how recently that SDK changed. Which means this isn't a one-time audit. The gap:

So it's something to monitor, not measure once. A library that scores 95 today can drop to 70 the week it ships v-next — then climb back as the next model generation learns it.

I'd rather you trust the number than oversell it:

None of that breaks the signal — it just scopes it. "Does the model reach for API surface that still exists?" is a real, useful question, and the compiler answers it without opinion.

Adding an SDK is about an afternoon: install the package, add a tsconfig

, write a tasks file, register it. The harness handles generation, type-checking, and scoring.

I'm looking for the next SDKs to score. What library have you watched an AI assistant get wrong since its last major? Tell me and I'll run it.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @sdkproof 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-coding-agents-sti…] indexed:0 read:4min 2026-07-20 ·