# We started running every MCP server we grade. Here's what 20 popular ones actually did.

> Source: <https://dev.to/agentavow/we-started-running-every-mcp-server-we-grade-heres-what-20-popular-ones-actually-did-53ik>
> Published: 2026-10-02 20:34:26+00:00

Both big AI platforms check an MCP server the same way before listing it: they verify the domain, read the self-declared annotations (`readOnlyHint`, `destructiveHint`), and scan the policy text. Then they contain the tool at runtime. Nobody checks whether the tool *behaves* the way it declares.

So we built that check into [AgentAvow](https://agentavow.com), and this post is about what happened when we pointed it at 20 well-known servers.

For any MCP server published as an npm or PyPI package (and now GitHub repos and OpenClaw skills), the first scan kicks off a run in a fresh gVisor container with a read-only root, no capabilities, one CPU, and a hard time limit:

`BehavioralObservation` anyone can verify offline.
The grading is a short list of fixed rules: egress to a host that is neither the package registry, the tool's own vendor, nor anything it declared; a tool that claims `readOnlyHint: true` and then writes files; a planted credential showing up in outbound traffic; a secret echoed back in a tool result.

The sandbox result feeds the trust score by public rules, so the number stays recomputable:

| Observed | Effect | 
|---|---|
| A planted credential left the sandbox, or any critical behavioral finding | capped at 45 | 
| A high finding (undeclared egress, a read-only tool that writes) | −10, capped at 70 | 
| A clean, full exercise | +3 | 
| Didn't start, still running, or unsigned | no effect | 

The attestation records the observation's hash, the static score, and the delta, so you can check the arithmetic yourself.

We ran the 20 most-used MCP servers we could find on npm and PyPI.

That last category is the one people miss. Nothing malicious, nothing a static scan would flag, and your agent's tool calls are being reported to a company you've never heard of.

Validating against real servers found three false-positive patterns in our own grading, which is the part of this work I'd trust least if it weren't written down:

`AWS_REGION`.` 169.254.169.254`) was flagged as undeclared egress. It's now its own low-severity note: normal for cloud SDKs, worth knowing for anything else.
A clean sandbox run means the tool did nothing bad *under these conditions*. Code that waits for a date, checks for a real credential, or detects the sandbox will look clean. That is why the result is shown as an observation rather than a proof, why the static scan still looks for sandbox-probing code, and why we re-scan watched tools and alert on behavioral drift.

`fail_on_behavioral: true`.
The rules, the surfaces, and what each finding means are on the [behavioral sandbox page](https://agentavow.com/docs/behavioral-sandbox).
