{"slug": "we-started-running-every-mcp-server-we-grade-here-s-what-20-popular-ones-did", "title": "We started running every MCP server we grade. Here's what 20 popular ones actually did.", "summary": "AgentAvow built a behavioral sandbox that runs MCP servers in a fresh gVisor container with a read-only root, no capabilities, one CPU and a hard time limit, then grades them against fixed rules for undeclared egress, read-only tools that write files, planted credentials leaving the sandbox and secrets echoed in tool results. Running the 20 most-used MCP servers on npm and PyPI surfaced undeclared egress, including tools reporting agent tool calls to unknown third parties, and the team documented three false-positive patterns in its own grading, including AWS SDK metadata-endpoint traffic and clean runs that only reflect benign sandbox conditions.", "body_md": "Both big AI platforms check an MCP server the same way before listing it: they verify the domain, read the self-declared annotations (`readOnlyHint`, `destructiveHint`), and scan the policy text. Then they contain the tool at runtime. Nobody checks whether the tool *behaves* the way it declares.\n\nSo we built that check into [AgentAvow](https://agentavow.com), and this post is about what happened when we pointed it at 20 well-known servers.\n\nFor any MCP server published as an npm or PyPI package (and now GitHub repos and OpenClaw skills), the first scan kicks off a run in a fresh gVisor container with a read-only root, no capabilities, one CPU, and a hard time limit:\n\n`BehavioralObservation` anyone can verify offline.\nThe grading is a short list of fixed rules: egress to a host that is neither the package registry, the tool's own vendor, nor anything it declared; a tool that claims `readOnlyHint: true` and then writes files; a planted credential showing up in outbound traffic; a secret echoed back in a tool result.\n\nThe sandbox result feeds the trust score by public rules, so the number stays recomputable:\n\n| Observed | Effect | \n|---|---|\n| A planted credential left the sandbox, or any critical behavioral finding | capped at 45 | \n| A high finding (undeclared egress, a read-only tool that writes) | −10, capped at 70 | \n| A clean, full exercise | +3 | \n| Didn't start, still running, or unsigned | no effect | \n\nThe attestation records the observation's hash, the static score, and the delta, so you can check the arithmetic yourself.\n\nWe ran the 20 most-used MCP servers we could find on npm and PyPI.\n\nThat last category is the one people miss. Nothing malicious, nothing a static scan would flag, and your agent's tool calls are being reported to a company you've never heard of.\n\nValidating against real servers found three false-positive patterns in our own grading, which is the part of this work I'd trust least if it weren't written down:\n\n`AWS_REGION`.` 169.254.169.254`) was flagged as undeclared egress. It's now its own low-severity note: normal for cloud SDKs, worth knowing for anything else.\nA clean sandbox run means the tool did nothing bad *under these conditions*. Code that waits for a date, checks for a real credential, or detects the sandbox will look clean. That is why the result is shown as an observation rather than a proof, why the static scan still looks for sandbox-probing code, and why we re-scan watched tools and alert on behavioral drift.\n\n`fail_on_behavioral: true`.\nThe rules, the surfaces, and what each finding means are on the [behavioral sandbox page](https://agentavow.com/docs/behavioral-sandbox).", "url": "https://wpnews.pro/news/we-started-running-every-mcp-server-we-grade-here-s-what-20-popular-ones-did", "canonical_source": "https://dev.to/agentavow/we-started-running-every-mcp-server-we-grade-heres-what-20-popular-ones-actually-did-53ik", "published_at": "2026-10-02 20:34:26+00:00", "updated_at": "2026-10-02 20:37:11.806309+00:00", "lang": "en", "topics": ["ai-agents", "agent-protocols", "ai-safety", "developer-tools", "ai-tools"], "entities": ["AgentAvow", "MCP", "npm", "PyPI", "GitHub", "OpenClaw", "gVisor", "AWS"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/we-started-running-every-mcp-server-we-grade-here-s-what-20-popular-ones-did", "markdown": "https://wpnews.pro/news/we-started-running-every-mcp-server-we-grade-here-s-what-20-popular-ones-did.md", "text": "https://wpnews.pro/news/we-started-running-every-mcp-server-we-grade-here-s-what-20-popular-ones-did.txt", "jsonld": "https://wpnews.pro/news/we-started-running-every-mcp-server-we-grade-here-s-what-20-popular-ones-did.jsonld"}}