Both big AI platforms check an MCP server the same way before listing it: they verify the domain, read the self-declared annotations (readOnlyHint, destructiveHint), and scan the policy text. Then they contain the tool at runtime. Nobody checks whether the tool behaves the way it declares.
So we built that check into AgentAvow, and this post is about what happened when we pointed it at 20 well-known servers.
For any MCP server published as an npm or PyPI package (and now GitHub repos and OpenClaw skills), the first scan kicks off a run in a fresh gVisor container with a read-only root, no capabilities, one CPU, and a hard time limit:
BehavioralObservation anyone can verify offline.
The grading is a short list of fixed rules: egress to a host that is neither the package registry, the tool's own vendor, nor anything it declared; a tool that claims readOnlyHint: true and then writes files; a planted credential showing up in outbound traffic; a secret echoed back in a tool result.
The sandbox result feeds the trust score by public rules, so the number stays recomputable:
| Observed | Effect |
|---|---|
| A planted credential left the sandbox, or any critical behavioral finding | capped at 45 |
| A high finding (undeclared egress, a read-only tool that writes) | −10, capped at 70 |
| A clean, full exercise | +3 |
| Didn't start, still running, or unsigned | no effect |
The attestation records the observation's hash, the static score, and the delta, so you can check the arithmetic yourself.
We ran the 20 most-used MCP servers we could find on npm and PyPI.
That last category is the one people miss. Nothing malicious, nothing a static scan would flag, and your agent's tool calls are being reported to a company you've never heard of.
Validating against real servers found three false-positive patterns in our own grading, which is the part of this work I'd trust least if it weren't written down:
AWS_REGION. 169.254.169.254) was flagged as undeclared egress. It's now its own low-severity note: normal for cloud SDKs, worth knowing for anything else.
A clean sandbox run means the tool did nothing bad under these conditions. Code that waits for a date, checks for a real credential, or detects the sandbox will look clean. That is why the result is shown as an observation rather than a proof, why the static scan still looks for sandbox-probing code, and why we re-scan watched tools and alert on behavioral drift.
fail_on_behavioral: true.
The rules, the surfaces, and what each finding means are on the behavioral sandbox page.