Your MCP server changed. Its version didn't. Here's how to catch it. A developer built @sbissoli/mcp-surface, an npm tool that locks an MCP server's declared and no-token surfaces into a surface.lock.json file and fails tests, deploys or publishes when the surface changes without a version bump. Running the check backwards across 142 published versions of seven servers surfaced drift where a registry entry still said 0.1.0 while the live server had grown from four read-only tools to six, two of which write on the user's behalf. A version number is a promise: nothing a client depends on changed without me saying so . For MCP servers, almost nothing checks that promise. The copy of a server in the MCP Registry carries a name, a version, the packages and the remotes. It carries no surface: no tools, no instructions, nothing about what answers without a token. So anyone comparing the registry with the live server can compare exactly one thing, the version. That comparison is only worth something if every surface change bumps the version. This post is about turning that "if" into a failing test, and then about what happened when I ran the test backwards over the whole history of my seven servers: 142 published versions. It started in the comments of my previous post https://dev.to/sidneybissoli/building-an-mcp-server-for-financial-data-lessons-learned-24gi , in a thread with @yahhi https://dev.to/yahhi . The registry listing and the live worklore server both said 0.1.0 , so a version diff passed cleanly. But, as @yahhi found https://dev.to/yahhi/comment/3g5m4 , the server had moved on: tools/list now answered without a token, and the auth rules had moved into the instructions. The drift was one level deeper than the version. Next came a working check https://dev.to/yahhi/comment/3g607 , and a replay of the surface from the day of the listing against the current code. The registry said 0.1.0 for a server with four read-only tools. The server behind that number had six, two of which write on the user's behalf. I said I'd build the same thing for my servers. It's now live on all seven, and packaged as @sbissoli/mcp-surface https://www.npmjs.com/package/@sbissoli/mcp-surface . The lock file, surface.lock.json , sits at the root of each server, next to package.json . It has two sections. Each stores the version it was locked under and a sha256 of its content. initialize result instructions and capabilities included, serverInfo without the version , plus tools/list , resources/list , resources/templates/list and prompts/list . Keys are sorted and lists are ordered by name or URI, so the hash only moves when the surface does. A method the server doesn't serve is recorded as null , not : "no prompts" and "prompts/list isn't served" are different surfaces. The declared surface is captured from the server factory in-process. The no-token section is measured at the HTTP edge, by calling the Worker's fetch with and without the key. The rule is short: The script that rewrites the lock follows the same rule: it refuses to lock a new surface under the old version. You can't "fix" a red test by regenerating the lock. You bump the version, re-lock, and commit the lock with the bump: npm version minor --no-git-tag-version npm run surface:lock build + run the surface tests in write mode The deploy runs npm test before wrangler , and the npm publish runs it before npm publish . So a surface change without a version change can't be deployed or published. One side effect, and it's on purpose: an SDK upgrade that changes capabilities also turns the lock red. The client sees a different surface, so the version should say so. Tests run against the code in the repository. The client talks to whatever is actually deployed. Usually those are the same thing. Wrangler configs, build steps, environment variables and caches are where they stop being the same. So every deploy now ends with one more step that queries the live endpoint and compares it with the lock: npx mcp-surface verificar https://