Your coding agent shouldn't run pytest A developer has built verdict, an MCP server that gives coding agents structured, sandboxed test feedback, replacing raw pytest output with concise typed verdicts. The tool runs tests in ephemeral containers, fingerprints failures for history tracking, and flags preexisting breakage, reducing context token usage from ~40k to ~400 tokens. The project is open source and has been integrated with Claude Code, with known gaps documented in SECURITY.md. First post in a build-in-public series about verdict, an MCP server that gives coding agents structured, sandboxed test feedback. Watch a coding agent work and you'll see it run pytest in your shell, unsandboxed, and then push 40,000 tokens of raw output through its context window to answer one question: did my change break anything? That's three problems in one command: verdict is an MCP server that replaces the pytest shell-out with four tools: | tool | what it returns | |---|---| verify scope? | impact-selected tests, run in an ephemeral container, as a ~400-token typed verdict | explain failure check id | the full traceback - only on demand | history fingerprint | first seen / last seen / times seen for a failure | run checks "ruff","mypy" | lint & type checks, same verdict shape | ▶️ Watch the 30-second demo https://github.com/Dgotlieb/verdict-mcp readme - Claude Code fixing a bug with verdict verifying in a container. 1. Verdicts, not output. verify returns typed JSON: counts, per-failure message + location, and nothing else. Full tracebacks live behind explain failure . The whole verdict for a real failing run is ~400 tokens - the raw pytest output it replaces was ~40k. The design rule in the repo is blunt: nothing bulky rides in the summary, ever. 2. Fingerprints give failures identity. Every failure is hashed from its normalized signature - volatile tokens addresses, tmp paths, ids, durations collapsed first. Same logical failure ⇒ same fingerprint, across runs and refactors. Fingerprints are what make the third idea possible: 3. History answers "was it me?" verdict keeps a small SQLite db per project. Every failure in a verdict carries preexisting: true|false - this exact failure was known before your change vs. never seen it, it's yours . In the demo session that flag is the difference between an agent politely ignoring long-standing breakage and an agent burning a session "fixing" it. Checks run in an ephemeral container podman preferred, docker fallback, auto-detected - with liveness checks, because a podman binary with a stopped VM is worse than no podman at all . The worktree is mounted read-only at /src , copied to a writable /work inside, and the check run gets --network=none . Configured setup cmd runs with network before the check; a prebuilt image is the tighter posture. No engine? An explicit prefer = "local" fallback still runs against a temp copy - and if verdict ever runs somewhere other than where you configured, it says so in the verdict runner note . No silent degradation is a design rule. Known gaps are written down in SECURITY.md rather than hand-waved: no resource limits yet, setup cmd is network-open by design. I wired verdict into Claude Code and asked it to fix a bug. Two things surfaced in the first hour that the 26-test suite had missed: sys.path hack and an "honest fallback is acceptable" escape hatch. Deleted both, fixed for real. pytest --json-report-omit takes nargs='+' … and the test paths came right after it. pytest swallowed them as omit values. Ten characters of argument reordering.Both bugs made the headline feature a no-op while all tests were green. The suite now has the tests that would have caught them - written by breaking the real thing first. Cloned python-dotenv https://github.com/theskumar/python-dotenv 255 tests , dropped in a two-line verdict.toml , no other changes: pip install variables.py → impact selection runs 184 tests, skips ~70, and says exactly how approximate the selection is in selection note That 146s is the next mountain: env setup needs caching image reuse keyed on deps, result cache keyed on tree hash before verify feels instant. That's v0.2, and it's written down as v0.2 - scope discipline is also a feature. { "mcpServers": { "verdict": { "command": "uvx", "args": "verdict-mcp" , "env": { "VERDICT PROJECT": "." } } } } pip install verdict-mcp / uvx verdict-mcp . Apache-2.0. Demo video and the full threat model in the repo https://github.com/Dgotlieb/verdict-mcp .