We started running every MCP server we grade. Here's what 20 popular ones actually did. AgentAvow built a behavioral sandbox that runs MCP servers in a fresh gVisor container with a read-only root, no capabilities, one CPU and a hard time limit, then grades them against fixed rules for undeclared egress, read-only tools that write files, planted credentials leaving the sandbox and secrets echoed in tool results. Running the 20 most-used MCP servers on npm and PyPI surfaced undeclared egress, including tools reporting agent tool calls to unknown third parties, and the team documented three false-positive patterns in its own grading, including AWS SDK metadata-endpoint traffic and clean runs that only reflect benign sandbox conditions. Both big AI platforms check an MCP server the same way before listing it: they verify the domain, read the self-declared annotations readOnlyHint , destructiveHint , and scan the policy text. Then they contain the tool at runtime. Nobody checks whether the tool behaves the way it declares. So we built that check into AgentAvow https://agentavow.com , and this post is about what happened when we pointed it at 20 well-known servers. For any MCP server published as an npm or PyPI package and now GitHub repos and OpenClaw skills , the first scan kicks off a run in a fresh gVisor container with a read-only root, no capabilities, one CPU, and a hard time limit: BehavioralObservation anyone can verify offline. The grading is a short list of fixed rules: egress to a host that is neither the package registry, the tool's own vendor, nor anything it declared; a tool that claims readOnlyHint: true and then writes files; a planted credential showing up in outbound traffic; a secret echoed back in a tool result. The sandbox result feeds the trust score by public rules, so the number stays recomputable: | Observed | Effect | |---|---| | A planted credential left the sandbox, or any critical behavioral finding | capped at 45 | | A high finding undeclared egress, a read-only tool that writes | −10, capped at 70 | | A clean, full exercise | +3 | | Didn't start, still running, or unsigned | no effect | The attestation records the observation's hash, the static score, and the delta, so you can check the arithmetic yourself. We ran the 20 most-used MCP servers we could find on npm and PyPI. That last category is the one people miss. Nothing malicious, nothing a static scan would flag, and your agent's tool calls are being reported to a company you've never heard of. Validating against real servers found three false-positive patterns in our own grading, which is the part of this work I'd trust least if it weren't written down: AWS REGION . 169.254.169.254 was flagged as undeclared egress. It's now its own low-severity note: normal for cloud SDKs, worth knowing for anything else. A clean sandbox run means the tool did nothing bad under these conditions . Code that waits for a date, checks for a real credential, or detects the sandbox will look clean. That is why the result is shown as an observation rather than a proof, why the static scan still looks for sandbox-probing code, and why we re-scan watched tools and alert on behavioral drift. fail on behavioral: true . The rules, the surfaces, and what each finding means are on the behavioral sandbox page https://agentavow.com/docs/behavioral-sandbox .