{"slug": "i-built-tools-to-verify-what-my-ai-coding-agent-actually-did", "title": "I built tools to verify what my AI coding agent actually did", "summary": "A QA engineer who runs the Devin coding agent daily built a suite of twenty local-first tools that verify an agent's claims against its own on-disk telemetry rather than its narrative, after a session reported \"fixed, all green\" while the file on disk disagreed. The stack spans a runtime spec, a QA audit pack that checks diffs, tests, commits and pushes against tool_call_state, an eval suite, an ACP-based policy gate, and a calibration judge that improved expected calibration error from 0.170 to 0.071 on real session data. All tools run on a RAM-constrained laptop with no telemetry leaving the machine, and the judge is exposed as an MCP server via `poordjaevin serve`.", "body_md": "AI coding agents are great at one thing that isn't writing code: *asserting*. \"Tests pass.\" \"The file was updated.\" \"I pushed the fix.\" And if you've run agents for real work, you know these claims are sometimes... optimistic.\n\nI'm a QA engineer who runs [Devin](https://devin.ai) daily. At some point I got tired of manually checking whether the agent's claims matched reality, so I did what a QA engineer does: I built a test harness. Then an eval suite. Then a policy layer. Then a decision layer. Twenty local-first tools later, the whole stack runs on a RAM-constrained laptop with zero telemetry leaving the machine.\n\nThis is the short version of what exists, why, and what I learned.\n\nThe trigger was simple. An agent session reported \"fixed, all green\" — and the file on disk disagreed. No malice; agents report from their own narrative, not from ground truth. Classic QA problem: the *claim* and the *state* are different artifacts, and only one of them is evidence.\n\nThe insight that made everything else possible: agent CLIs already write structured telemetry locally — session files, tool-call state, token usage. The evidence was sitting on disk the whole time. I just had to stop trusting the story and start reading the log.\n\nThe tools compose into a sequence: understand → verify → measure → control → judge.\n\n**[devin-internals-spec](https://github.com/Icaro0310/devin-internals-spec)** — *understand.* Before you can trust tool output you need to know the contracts: the file formats, exit codes, and behavioral rules of the runtime itself. This is the spec layer — everything downstream reads against it.\n\n**[devin-qa-pack](https://github.com/Icaro0310/devin-qa-pack)** — *verify.* Runs a QA audit over actual agent work: file diffs present, tests run, commits exist, pushes landed, verification commands executed. Claims are checked against `tool_call_state`, not against the agent's narrative. 47 tests, CI on Ubuntu + Windows. This is the tool that started it all.\n\n**[devin-evals](https://github.com/Icaro0310/devin-evals)** — *measure.* Once verification works, you can ask the harder question: how good is the agent on *this* kind of task? Golden tasks, rubric scoring, regression tracking across sessions.\n\n**[devin-bridge](https://github.com/Icaro0310/devin-bridge)** — *control.* A policy gate between intent and execution — ACP-based control with a `--devin-only` mode that enforces \"this session does exactly what it was scoped to do.\"\n\n**[poordjaevin](https://github.com/Icaro0310/poordjaevin)** — *judge.* The piece I'm proudest of. Takes a task description and produces a calibrated confidence score: should I trust this delegation? The calibration story is the interesting part — the judge went from **ECE 0.170 to 0.071** through iterative refinement on real session data. (ECE = expected calibration error: when the judge says \"80% confident,\" it should be right ~80% of the time. Most confidence scores don't do this. Now mine roughly does.) It's also a real MCP server — `poordjaevin serve` — so agents can consult it mid-flight.\n\nEvery tool follows the same contract: read local agent state, write local artifacts, expose a stable CLI, degrade gracefully when optional infrastructure is absent. No required daemons, no SaaS dashboard, no telemetry.\n\nConcretely:\n\n`devin-history` — cross-session memory: searchable SQLite over all past sessions (7+ commands of grep-able agent archaeology)`devin-metrics` — telemetry aggregation for quality signals over time`devin-memory` + `devin-search` + `devin-graph` — memory store, retrieval, and a knowledge graph over sessions, projects and decisions`devin-doctor` — environment diagnostics across Windows/Linux`devin-backup` — snapshot + verify + restore of agent state`devin-janitor` + `devin-redact` — cleanup and secret-redaction so state can move safely`devin-office` — the fun one: a live circuit-board dashboard that renders real sessions, subagents and tool calls from the local store. Purely visual, read-only, zero telemetry — and it makes for a great demo GIF.\n**1. The agent's own telemetry is an untapped QA datasource.** Session files and tool-call state are structured evidence most people ignore. Reading them turns \"trust the agent\" into \"verify the agent\" — and verification is what makes delegation safe at scale.\n\n**2. Calibration > confidence.** A judge that's confidently wrong is worse than no judge. Iterating on ECE (0.170 → 0.071) took real labeled outcomes, not prompt tweaks. If you build any kind of AI decision layer, measure calibration explicitly.\n\n**3. Devin-only mode matters more than integrations.** Every tool works standalone with just the agent's CLI present. Obsidian, Slack, MCP — all optional. The ecosystem must survive a locked-down corporate box with nothing but Devin installed, because that's where a lot of real work happens.\n\nEverything is MIT-licensed, cross-platform, and documented in English + PT-BR:\n\n`devin-qa-pack` (the verifier) or `poordjaevin` (the calibrated judge — `pip install poordjaevin` / `uv tool install poordjaevin`)\nHonest question for the comments: if you run AI agents on real work — Copilot, Devin, Cursor, Claude Code — **how do you verify their claims today?** Manual spot-checks? CI gates? Nothing? I suspect \"nothing\" is the most common answer, and I built this stack partly because it scared me that it was mine.", "url": "https://wpnews.pro/news/i-built-tools-to-verify-what-my-ai-coding-agent-actually-did", "canonical_source": "https://dev.to/icaro0310/i-built-tools-to-verify-what-my-ai-coding-agent-actually-did-2a6f", "published_at": "2026-10-06 17:12:11+00:00", "updated_at": "2026-10-06 17:18:58.360517+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "agent-protocols", "mlops"], "entities": ["Devin", "devin-qa-pack", "devin-evals", "devin-bridge", "poordjaevin", "devin-internals-spec", "MCP", "GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-built-tools-to-verify-what-my-ai-coding-agent-actually-did", "markdown": "https://wpnews.pro/news/i-built-tools-to-verify-what-my-ai-coding-agent-actually-did.md", "text": "https://wpnews.pro/news/i-built-tools-to-verify-what-my-ai-coding-agent-actually-did.txt", "jsonld": "https://wpnews.pro/news/i-built-tools-to-verify-what-my-ai-coding-agent-actually-did.jsonld"}}