I built tools to verify what my AI coding agent actually did A QA engineer who runs the Devin coding agent daily built a suite of twenty local-first tools that verify an agent's claims against its own on-disk telemetry rather than its narrative, after a session reported "fixed, all green" while the file on disk disagreed. The stack spans a runtime spec, a QA audit pack that checks diffs, tests, commits and pushes against tool_call_state, an eval suite, an ACP-based policy gate, and a calibration judge that improved expected calibration error from 0.170 to 0.071 on real session data. All tools run on a RAM-constrained laptop with no telemetry leaving the machine, and the judge is exposed as an MCP server via `poordjaevin serve`. AI coding agents are great at one thing that isn't writing code: asserting . "Tests pass." "The file was updated." "I pushed the fix." And if you've run agents for real work, you know these claims are sometimes... optimistic. I'm a QA engineer who runs Devin https://devin.ai daily. At some point I got tired of manually checking whether the agent's claims matched reality, so I did what a QA engineer does: I built a test harness. Then an eval suite. Then a policy layer. Then a decision layer. Twenty local-first tools later, the whole stack runs on a RAM-constrained laptop with zero telemetry leaving the machine. This is the short version of what exists, why, and what I learned. The trigger was simple. An agent session reported "fixed, all green" — and the file on disk disagreed. No malice; agents report from their own narrative, not from ground truth. Classic QA problem: the claim and the state are different artifacts, and only one of them is evidence. The insight that made everything else possible: agent CLIs already write structured telemetry locally — session files, tool-call state, token usage. The evidence was sitting on disk the whole time. I just had to stop trusting the story and start reading the log. The tools compose into a sequence: understand → verify → measure → control → judge. devin-internals-spec https://github.com/Icaro0310/devin-internals-spec — understand. Before you can trust tool output you need to know the contracts: the file formats, exit codes, and behavioral rules of the runtime itself. This is the spec layer — everything downstream reads against it. devin-qa-pack https://github.com/Icaro0310/devin-qa-pack — verify. Runs a QA audit over actual agent work: file diffs present, tests run, commits exist, pushes landed, verification commands executed. Claims are checked against tool call state , not against the agent's narrative. 47 tests, CI on Ubuntu + Windows. This is the tool that started it all. devin-evals https://github.com/Icaro0310/devin-evals — measure. Once verification works, you can ask the harder question: how good is the agent on this kind of task? Golden tasks, rubric scoring, regression tracking across sessions. devin-bridge https://github.com/Icaro0310/devin-bridge — control. A policy gate between intent and execution — ACP-based control with a --devin-only mode that enforces "this session does exactly what it was scoped to do." poordjaevin https://github.com/Icaro0310/poordjaevin — judge. The piece I'm proudest of. Takes a task description and produces a calibrated confidence score: should I trust this delegation? The calibration story is the interesting part — the judge went from ECE 0.170 to 0.071 through iterative refinement on real session data. ECE = expected calibration error: when the judge says "80% confident," it should be right ~80% of the time. Most confidence scores don't do this. Now mine roughly does. It's also a real MCP server — poordjaevin serve — so agents can consult it mid-flight. Every tool follows the same contract: read local agent state, write local artifacts, expose a stable CLI, degrade gracefully when optional infrastructure is absent. No required daemons, no SaaS dashboard, no telemetry. Concretely: devin-history — cross-session memory: searchable SQLite over all past sessions 7+ commands of grep-able agent archaeology devin-metrics — telemetry aggregation for quality signals over time devin-memory + devin-search + devin-graph — memory store, retrieval, and a knowledge graph over sessions, projects and decisions devin-doctor — environment diagnostics across Windows/Linux devin-backup — snapshot + verify + restore of agent state devin-janitor + devin-redact — cleanup and secret-redaction so state can move safely devin-office — the fun one: a live circuit-board dashboard that renders real sessions, subagents and tool calls from the local store. Purely visual, read-only, zero telemetry — and it makes for a great demo GIF. 1. The agent's own telemetry is an untapped QA datasource. Session files and tool-call state are structured evidence most people ignore. Reading them turns "trust the agent" into "verify the agent" — and verification is what makes delegation safe at scale. 2. Calibration confidence. A judge that's confidently wrong is worse than no judge. Iterating on ECE 0.170 → 0.071 took real labeled outcomes, not prompt tweaks. If you build any kind of AI decision layer, measure calibration explicitly. 3. Devin-only mode matters more than integrations. Every tool works standalone with just the agent's CLI present. Obsidian, Slack, MCP — all optional. The ecosystem must survive a locked-down corporate box with nothing but Devin installed, because that's where a lot of real work happens. Everything is MIT-licensed, cross-platform, and documented in English + PT-BR: devin-qa-pack the verifier or poordjaevin the calibrated judge — pip install poordjaevin / uv tool install poordjaevin Honest question for the comments: if you run AI agents on real work — Copilot, Devin, Cursor, Claude Code — how do you verify their claims today? Manual spot-checks? CI gates? Nothing? I suspect "nothing" is the most common answer, and I built this stack partly because it scared me that it was mine.