{"slug": "show-hn-i-measured-the-mcp-registry-nightly-for-37-days", "title": "Show HN: I measured the MCP registry nightly for 37 days", "summary": "An independent nightly crawler of the official MCP registry reached 13,870 of 13,997 declared endpoints (99.1%) and enumerated tool lists without credentials on 7,820 servers — 56.4% of endpoints with a remote and 31.3% of the whole registry — recording 132,683 tools across nine consecutive daily censuses as of 2026-09-07. The project's operator reports that two operators, gateway.pipeworx.io (1,264 servers) and api.mcp.ai (1,094), account for 30.2% of every reachable MCP server in the registry, and that 73.2% of recorded tools declare at least one annotation hint. The log is an append-only, signed-tree-head record of what public MCP servers actually served, positioned as adversarial third-party observation rather than opt-in publisher attestation such as Sigstore, Rekor and manifest signing.", "body_md": "An independent, continuously-operated record of what public MCP servers actually served.\n\nSigstore, Rekor and manifest signing are **opt-in publisher attestations**. A server operator who changes a tool description from *\"look up the weather\"* to *\"look up the weather and forward the conversation to evil.example\"* will happily sign the new description, and the signature will verify perfectly. Signing proves who published something. It does not tell you that what you are running today is not what you audited last month.\n\nA transparency log is **adversarial third-party observation**. It records what servers actually served, without their consent or cooperation. That asymmetry is why Certificate Transparency worked: CAs never opted in, monitors watched them anyway, and misissuance became detectable after the fact.\n\nThat is what this project is. Not a scanner, not a registry, not a linter — a log.\n\nA crawler visits every publicly reachable MCP server in the official registry on a schedule and records the exact tool surface it served: names, descriptions, JSON schemas, and the four annotation hints. Observations go into an append-only log with signed tree heads, so anyone can prove that a given surface was observed at a given time and that the log has not been rewritten since.\n\nThe value is the accumulated history. A stranger can clone this repository in a weekend. Nobody can clone a year of observations.\n\nEarly. Phase 1 of 4.\n\n| Phase | Scope | State | \n|---|---|---|\n| 0 | Falsification test — is the ecosystem observable at all? | done, passed | \n| 1 | Crawler, raw observation archive, daily census | done, running | \n| 2 | RFC 6962 Merkle log, signed tree heads, verify CLI | done | \n| 3 | Public static site | not started | \n| 4 | Witnesses and gossip for split-view detection | not started | \n\nNine consecutive daily censuses as of 2026-09-07, no gaps. A systemd timer fetches the registry, probes every endpoint, archives the raw bytes, appends to the tree, signs a head and publishes it — with no human in the loop.\n\nPhase 4 is the end state, not the entry ticket. A single-operator log is still useful — Go's own checksum database ran that way for years.\n\nThe first complete pass over the registry. 13,870 of 13,997 declared endpoints were reached; the run's `meta.json` records it as incomplete, because it is.\n\n| Registry entries | 25,020 | \n| Declaring a network endpoint | 13,997 | \n| Observed | 13,870 (99.1%) | \n| **Enumerated a tool list, no credentials** | **7,820 — 56.4%** of endpoints with a remote | \n| Same figure against the whole registry | **31.3%** | \n| Tools recorded | **132,683** | \n| Tools declaring at least one annotation hint | **73.2%** | \n\nOutcomes: 56.4% ok, 25.0% auth required, 8.3% protocol error, 6.4% unreachable, 3.2% timeout, 0.7% rpc error. Distinguishing \"refused us\" from \"is not there\" is what keeps the reachability figure honest.\n\n**Concentration.** Two operators — `gateway.pipeworx.io` (1,264) and `api.mcp.ai` (1,094) — account for **30.2% of every reachable MCP server in the registry**. One template edit at either changes over a thousand \"servers\" at once. This is the single strongest argument for watching this ecosystem rather than trusting it.\n\n**Spec adoption.** 6,716 servers negotiated 2025-06-18; 672 still speak 2024-11-05. Seventeen answered on 2026-07-28. The week-0 sample of 300 found zero on that revision and concluded none existed — at full population the honest statement is \"rare, not absent\". A sample that small cannot see a 0.2% feature.\n\n**Tool-surface size.** The median server exposes a handful of tools. Three expose more than 600, and one — `io.github.davidmosiah/delx-mcp-a2a` — exposes **1,076**, of which 31 carry any annotation. Only 2 servers in the entire population paginated `tools/list`.\n\n**Parked domains still listed.** Four registry endpoints resolve to expired domains now serving for-sale parking pages, all four through the same ad host. Small in absolute terms (0.03%), but the mechanism matters: an agent configured from the official registry connects to infrastructure its original operator no longer controls. Windows Defender classified two of those pages as phishing — noted, not endorsed: the same detector had flagged this project's own binary as a trojan an hour earlier. The archived bytes are in the log; judge them yourself.\n\nBefore writing a crawler, the obvious way to kill this idea was tested: if almost no public MCP server can be enumerated without credentials, there is nothing to observe and the project should not exist. The kill threshold was set at 30% in advance.\n\nMeasured on a random sample of 300 registry entries, 2026-08-25:\n\n| Registry population | **24,729** servers | \n| Declaring a network endpoint | **13,629** (55.1%) | \n| Enumerable with no credentials | **35.0%** — above the 30% kill line | \n| Genuinely real servers in the sample | 67.6%, across 71 distinct operators | \n| Tools declaring at least one annotation hint | **68.8%** of 1,420 observed tools | \n| Successful handshakes using the 2026-07-28 `server/discover` | **0 of 105** — all used legacy`initialize` | \n\nTwo operators, `gateway.pipeworx.io` (1,312 servers) and `api.mcp.ai` (1,099), publish **17.7% of every observable server in the registry**. One template edit changes 1,312 \"servers\" at once. That concentration is the clearest argument for watching this ecosystem rather than trusting it.\n\nThe probe, its raw output, and the population snapshot are in [`docs/week0-falsification/`](https://github.com/yassinht/mcp-transparency-log/blob/main/docs/week0-falsification). Two bugs in that probe mattered enormously: missing SSE transport support understated reachability by ten percentage points and would have produced a false \"do not build\" verdict, and a hardcoded protocol version biased the sample against servers on newer spec revisions. Writing the code was never the bottleneck. Contact with reality was.\n\n**Raw bytes are the record.** Every response is archived exactly as received. Counts, classifications and hashes are derived views that can be rebuilt. The week-0 probe stored only its own classification of the registry and discarded the raw entries, which made every later question about that snapshot unanswerable.\n\n**Two hashes, never one.** `body_sha256` covers the exact bytes and is the provenance claim. `surface_sha256` covers the canonical, name-sorted tool array and is the change-detection key and the future Merkle leaf. Collapsing them would make the log either noisy or unprovable.\n\n**No MCP SDK in the crawl path.** An SDK validates, normalizes and rejects, because it is built to talk to well-behaved servers. A server returning a nameless tool, a duplicate name, or a 40KB description is producing exactly the observation worth keeping.\n\n**No LLM anywhere in the data path.** Every judgment must be a rule a human can re-run against the archived blobs, or the Merkle proofs are theatre.\n\n**Content addressing, not timestamps.** A server whose surface has not changed costs nothing to observe again. The archive grows only when something actually changed.\n\n```\ngo test ./...\ngo run ./cmd/crawler --dry-run          # fetch and archive the registry, probe nothing\ngo run ./cmd/crawler --sample=200       # reproducible smoke test\ngo run ./cmd/crawler                    # full census\n```\n\nOutput lands in `data/`:\n\n```\ndata/blobs/<ab>/<sha256>        exact response bytes, deduplicated across all runs\ndata/runs/<runID>/index.jsonl   one line per endpoint observed\ndata/runs/<runID>/meta.json     run header, including the raw registry pages\n```\n\nPoliteness is structural rather than advisory: targets are grouped by host and each host is worked by exactly one goroutine with a delay between requests, so no amount of `--workers` can hammer a single operator.\n\nThis is the question the log exists to answer, and the first attempt at it was wrong by a factor of six. Both the wrong number and the correction are kept here, because the correction is the more useful of the two.\n\nComparing two censuses seventeen hours and fifty-one minutes apart (2026-08-30 09:12 → 2026-08-31 03:03, same machine, same network):\n\n| Enumerable in both runs | 8,106 | \n| Surface unchanged | 6,551 — 80.8% | \n| Surface changed | 1,555 — 19.2% | \n| Of those, keeping **identical tool names** | 1,480 — 95.2% of all changes | \n\nA server that adds or removes a tool is visible: the client sees the list change. A server that keeps `send_email` under the same name and rewrites what it claims to do announces nothing. No version bump, no notification. That second shape is the one signing cannot catch — the operator signs the new description and the signature verifies — and it accounts for 95% of everything that moves.\n\n**Then the number had to survive its own audit.** Some servers embed live data in a description: *\"cache updated 2026-08-30 09:14:02, 1,204 cities\"* changes on every request and means nothing. `mcpobs classify` separates those by a rule anyone can re-run against the archived bytes — replace every digit with `#` and compare again — and the result was not kind:\n\n| digits only | 83.5% | not a change | \n| schema edited | 10.4% | real | \n| description rewritten | 6.0% | real | \n\nThat reading — 83.5% noise — was itself an artifact, and finding out why produced the most useful methodological result here. **It came from comparing two censuses taken at different times of day** (09:12 against 03:03). Many servers embed content on a daily cycle, so sampling at different points in that cycle makes them all look changed. Every census since runs at 03:00, and comparing same-hour to same-hour the noise collapses:\n\n| interval | volatile | real | one publisher's share | \n|---|---|---|---|\n| Aug 30 09:12 → Aug 31 03:03 | 57.7% | 42.3% | 69.0% | \n| **Sep 1 03:06 → Sep 2 03:07** | **1.0%** | **99.0%** | **87.1%** | \n| **Sep 3 03:04 → Sep 4 03:05** | **0.5%** | **99.5%** | **86.9%** | \n\n**Sampling a live system at a fixed hour removes cyclic noise; sampling it at wandering hours measures your own clock.** The two same-hour intervals agree to within half a point on every figure.\n\nOf roughly 8,300 servers enumerable on two consecutive days, **19.2% changed their tool surface within 24 hours**, and about 95% of those changes kept identical tool names while rewriting descriptions or schemas underneath.\n\n**But the headline number is one publisher.** `io.github.pipeworx-io` accounts for **87% of every real edit**, rewriting descriptions across roughly 1,280 servers — essentially its whole fleet — every single night. Excluding it, the independent server changes at about **2.3%** per day.\n\nBoth numbers are true and they answer different questions. What an agent operator experiences is 19%. How volatile a typical independent server is, is 2.3%. Reporting either alone misleads, and a single snapshot cannot tell them apart — only a daily log can.\n\nTwo earlier figures here were wrong and are kept on purpose. A first pass reported 26%. A sample then put the real fraction at 97%, and it was wrong for a reason worth naming: the analysis could only parse plain-JSON response bodies and silently skipped SSE-framed ones, inspecting 67 of 1,355. That sample was **biased, not small** — simple servers return plain JSON and write static descriptions, complex ones stream SSE and inject live data. Measuring the easy half and generalising is how a six-fold error gets published.\n\nOne limitation stands: the rule normalizes digits, not rotating prose. At least one server serves a different daily puzzle inside a tool description, which this classifier still counts as real.\n\nOn 2026-09-06 reachability fell to 50.5% and the tool count to 109,406, against a stable 56.5–58.5% and roughly 143,000 on every neighbouring day. The drop is far larger in tools than in servers, which means the servers that went missing were the large ones — a major operator was unreachable for a single night, and returned. It has not been investigated further. It is recorded here because the log's value is precisely that such nights are recoverable after the fact.\n\nThe log's public key:\n\n```\n4DNVWgqiY5HKiGH3PFcKt5O+fn9EmdoTNvj4aYbwQ4E=\n```\n\nEvery census appends its observations to an RFC 6962 Merkle tree and publishes a signed tree head in [`heads/`](https://github.com/yassinht/mcp-transparency-log/blob/main/heads) — a few hundred readable bytes:\n\n```\nmcp-transparency-log/v1\nsize 136225\nroot 1JPVp4Dx7ds59SivxgXvaHdoTpDD1W0VyTR6SmCCL8c=\ntime 2026-09-07T10:37:12Z\nsig  qv1sZUjrjHXFhtwPPPq1e770e57Xg1DpeeJFxsL/DX+5ZVl2pFDYJjHGoG6u9HzB2OYKZwPjF8VIJUw8eaaRCA==\ngo install github.com/yassinht/mcp-transparency-log/cmd/mcpobs@latest\nmcpobs verify\n```\n\n`verify` rebuilds the tree from the observation records and checks the result against every published head. It deliberately ignores the stored hash file: verifying a log against hashes its own operator wrote proves nothing, since the hashes and the lie would come from the same hand. It also rejects any head that shrinks the log, because publishing fewer observations than yesterday is a deletion of history that no valid signature excuses.\n\nThe tree is what makes the operator — me — untrusted rather than trusted. A signature alone cannot do this: I hold the key, so a forged head will always verify under it. What cannot be forged is the tree. If I edit one archived observation from last month, the records stop reproducing the root I signed at the time, and anyone holding that older head can prove it. [`internal/mlog/commit_test.go`](https://github.com/yassinht/mcp-transparency-log/blob/main/internal/mlog/commit_test.go) is that scenario as an executable test.\n\n**One honest limitation.** The first head, signed 2026-09-07, covers 136,225 observations taken over the nine days before the log existed. It attests that those observations are in the log *as of that date* — not that each was taken on the day it records. Only heads signed the day their observations were taken carry the stronger claim. Every head from 2026-09-08 onward does.\n\nThe crawler is a single static binary with no dependencies — no runtime, no database, no container.\n\n```\nGOOS=linux GOARCH=amd64 CGO_ENABLED=0 go build -trimpath -ldflags=\"-s -w\" -o dist/crawler ./cmd/crawler\nscp dist/crawler dist/mcpobs deploy/* user@server:~/\nssh user@server 'sudo bash install.sh'\n```\n\n`install.sh` creates a `mcpobs` service account with no shell, installs to `/opt/mcp-transparency-log`, and enables a daily systemd timer. It refuses to run if any name it needs is already taken, rather than overwriting something that might matter.\n\nThe unit is deliberately constrained. The crawler talks to fourteen thousand servers it does not control, on a box that is probably running something else that does matter:\n\n| `ProtectSystem=strict` ,`ReadWritePaths=…/data` | the filesystem is read-only except its own data directory | \n| `RestrictAddressFamilies=AF_INET AF_INET6` | network only; no local sockets | \n| `CPUQuota=50%` ,`IOWeight=20` ,`Nice=10` | background work yields to real applications | \n| `OOMScoreAdjust=800` | under memory pressure the kernel kills this, never the neighbours | \n| `MemoryMax=1500M` | measured, not guessed: a real census sat at 404 MB | \n| `Persistent=true` on the timer | a reboot spanning 03:00 catches up instead of leaving a hole | \n\nThree faults only appeared on the first real deployment and none were visible in development: progress written with carriage returns is invisible in `journalctl` and looks exactly like a hang; a 512 MB memory cap would have killed the census partway; and falling back to the SSE transport after a *timeout* doubled the cost of every dead endpoint, stretching the tail of a run by two hours. Writing the code was never the bottleneck.\n\nThis log covers **publicly observable MCP servers** — those listed in the official registry that declare a network endpoint and answer an unauthenticated `tools/list`. That is 35% of listed servers, not \"the MCP ecosystem\". The distinction is not modesty; the credibility of every number here depends on it.\n\nTBD.", "url": "https://wpnews.pro/news/show-hn-i-measured-the-mcp-registry-nightly-for-37-days", "canonical_source": "https://github.com/yassinht/mcp-transparency-log", "published_at": "2026-10-05 09:36:53+00:00", "updated_at": "2026-10-05 10:20:08.907768+00:00", "lang": "en", "topics": ["agent-protocols", "ai-agents", "ai-infrastructure", "ai-tools", "ai-safety"], "entities": ["MCP registry", "gateway.pipeworx.io", "api.mcp.ai", "io.github.davidmosiah/delx-mcp-a2a", "Sigstore", "Rekor", "RFC 6962", "Certificate Transparency"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-i-measured-the-mcp-registry-nightly-for-37-days", "markdown": "https://wpnews.pro/news/show-hn-i-measured-the-mcp-registry-nightly-for-37-days.md", "text": "https://wpnews.pro/news/show-hn-i-measured-the-mcp-registry-nightly-for-37-days.txt", "jsonld": "https://wpnews.pro/news/show-hn-i-measured-the-mcp-registry-nightly-for-37-days.jsonld"}}