cd /news/ai-agents/your-mcp-server-changed-its-version-… · home › topics › ai-agents › article
[ARTICLE · art-146826] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Your MCP server changed. Its version didn't. Here's how to catch it.

A developer built @sbissoli/mcp-surface, an npm tool that locks an MCP server's declared and no-token surfaces into a surface.lock.json file and fails tests, deploys or publishes when the surface changes without a version bump. Running the check backwards across 142 published versions of seven servers surfaced drift where a registry entry still said 0.1.0 while the live server had grown from four read-only tools to six, two of which write on the user's behalf.

by read6 min views14 publishedOct 7, 2026

A version number is a promise: nothing a client depends on changed without me saying so. For MCP servers, almost nothing checks that promise.

The copy of a server in the MCP Registry carries a name, a version, the packages and the remotes. It carries no surface: no tools, no instructions, nothing about what answers without a token. So anyone comparing the registry with the live server can compare exactly one thing, the version. That comparison is only worth something if every surface change bumps the version.

This post is about turning that "if" into a failing test, and then about what happened when I ran the test backwards over the whole history of my seven servers: 142 published versions.

It started in the comments of my previous post, in a thread with @yahhi. The registry listing and the live worklore server both said 0.1.0, so a version diff passed cleanly. But, as @yahhi found, the server had moved on: tools/list now answered without a token, and the auth rules had moved into the instructions. The drift was one level deeper than the version.

Next came a working check, and a replay of the surface from the day of the listing against the current code. The registry said 0.1.0 for a server with four read-only tools. The server behind that number had six, two of which write on the user's behalf.

I said I'd build the same thing for my servers. It's now live on all seven, and packaged as @sbissoli/mcp-surface.

The lock file, surface.lock.json, sits at the root of each server, next to package.json. It has two sections. Each stores the version it was locked under and a sha256 of its content.

initialize result (instructions and capabilities included, serverInfo without the version), plus tools/list, resources/list, resources/templates/list and prompts/list. Keys are sorted and lists are ordered by name or URI, so the hash only moves when the surface does. A method the server doesn't serve is recorded as null, not []: "no prompts" and "prompts/list isn't served" are different surfaces. The declared surface is captured from the server factory in-process. The no-token section is measured at the HTTP edge, by calling the Worker's fetch with and without the key.

The rule is short:

The script that rewrites the lock follows the same rule: it refuses to lock a new surface under the old version. You can't "fix" a red test by regenerating the lock. You bump the version, re-lock, and commit the lock with the bump:

npm version minor --no-git-tag-version
npm run surface:lock   # build + run the surface tests in write mode

The deploy runs npm test before wrangler, and the npm publish runs it before npm publish. So a surface change without a version change can't be deployed or published.

One side effect, and it's on purpose: an SDK upgrade that changes capabilities also turns the lock red. The client sees a different surface, so the version should say so.

Tests run against the code in the repository. The client talks to whatever is actually deployed. Usually those are the same thing. Wrangler configs, build steps, environment variables and caches are where they stop being the same.

So every deploy now ends with one more step that queries the live endpoint and compares it with the lock:

npx mcp-surface verificar https://<host>/mcp --tool <a tool that needs no network>

If what's being served isn't what was locked, the deploy job goes red after the fact, which is still much better than never.

A lock written today says nothing about the versions that came before it. For those there's a replay:

npx mcp-surface replay --url https://<host>/mcp

For every version published on npm, it installs the package in a temporary directory (--ignore-scripts), starts it over stdio and captures the surface with the same normalisation the lock uses. Then it compares each version with the previous one, lists every removal that didn't come with a major bump, and checks the live endpoint against the surface of the version its /status reports. The output is a Markdown table and a JSON file in baselines/, committed to the repo.

Across the seven servers that came to 142 versions. In all seven, the live server matches the version it reports. It also found three things I didn't know.

Two old releases of my Senate server, 1.1.0 and 1.1.2, crash on startup:

Error: Dynamic require of "events" is not supported

The esbuild bundle emitted ESM that still contained a CommonJS require for a Node built-in. Nobody had noticed, me included, because nobody installs a two-version-old release on purpose. But npm would have offered both as valid versions indefinitely. Any version diff would have called them fine forever, because their version numbers are perfectly well-formed. They're deprecated now.

The lesson: "the version string is right" and "the version runs" are different claims, and only one of them shows up in a registry.

My medical terminology server went from 1.0.2 to 1.1.x and six SNOMED CT tools disappeared from the default surface. They moved behind a licensing flag, off by default.

The CHANGELOG for that release even called it "technically breaking". It shipped as a minor anyway. There was a reason, since SNOMED access depends on a licence the user has to hold. But the client that was calling snomed_search yesterday doesn't read my reasoning. It gets "unknown tool".

The replay lists cases like this without judging them. Some removals have a good reason, and the CHANGELOG is where that reason lives. What matters is that the list exists, so the decision is made on purpose rather than discovered later by a user.

The first version of the replay flagged ilo-mcp-server 0.5.0 → 0.6.0, where a parameter became required, as a breaking change outside a major release. But under semver, 0.x is allowed to break in a minor. The line of compatibility in 0.x is the minor, not the major.

The rule is fixed now: in 0.x, breaks are compared within the same minor. I mention it because a tool that reports drift has to earn its alarms. One false alarm on day one teaches people to ignore the next real one.

This is the one I'd most like other people to avoid.

In one server (sih-br-mcp) the MCP backend sits behind an edge proxy, so the no-token test can't call the real handler in-process. My first test double for that backend answered result to every method. The lock was written from it, and it faithfully recorded "prompts/list answers" for a server that has no prompts.

All the tests passed, because the tests were checking the lock against the same lying double. The post-deploy check against the live endpoint is what caught it.

The fix has two parts. A double that answers everything is not a double of your server, so where the real handler can run in-process, run it. And the probe's sanity checks (a method that must answer, a method that must not) now run before anything is written to the lock, not after. A lock written from a broken probe is worse than no lock, because it makes a wrong surface look verified.

The package is MIT: @sbissoli/mcp-surface, source in mcp-br-commons. Fair warning: the API and CLI verbs are in Portuguese (travar = lock, verificar = verify), and the roadmap follows my own servers' needs.

Thanks to @yahhi for the thread that started all this. There is also a write-up of the other side of it, in Russian, on Habr. One comment thread changed how eight servers ship.

── more in #ai-agents 4 stories · sorted by recency
── more on @mcp registry 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/your-mcp-server-chan…] indexed:0 read:6min 2026-10-07 · —