cd /news/ai-agents/i-built-tools-to-verify-what-my-ai-c… Β· home β€Ί topics β€Ί ai-agents β€Ί article
[ARTICLE Β· art-146204] src=dev.to β†— pub= topic=ai-agents verified=true sentiment=↑ positive

I built tools to verify what my AI coding agent actually did

A QA engineer who runs the Devin coding agent daily built a suite of twenty local-first tools that verify an agent's claims against its own on-disk telemetry rather than its narrative, after a session reported "fixed, all green" while the file on disk disagreed. The stack spans a runtime spec, a QA audit pack that checks diffs, tests, commits and pushes against tool_call_state, an eval suite, an ACP-based policy gate, and a calibration judge that improved expected calibration error from 0.170 to 0.071 on real session data. All tools run on a RAM-constrained laptop with no telemetry leaving the machine, and the judge is exposed as an MCP server via `poordjaevin serve`.

read4 min views3 publishedOct 6, 2026

AI coding agents are great at one thing that isn't writing code: asserting. "Tests pass." "The file was updated." "I pushed the fix." And if you've run agents for real work, you know these claims are sometimes... optimistic.

I'm a QA engineer who runs Devin daily. At some point I got tired of manually checking whether the agent's claims matched reality, so I did what a QA engineer does: I built a test harness. Then an eval suite. Then a policy layer. Then a decision layer. Twenty local-first tools later, the whole stack runs on a RAM-constrained laptop with zero telemetry leaving the machine.

This is the short version of what exists, why, and what I learned.

The trigger was simple. An agent session reported "fixed, all green" β€” and the file on disk disagreed. No malice; agents report from their own narrative, not from ground truth. Classic QA problem: the claim and the state are different artifacts, and only one of them is evidence.

The insight that made everything else possible: agent CLIs already write structured telemetry locally β€” session files, tool-call state, token usage. The evidence was sitting on disk the whole time. I just had to stop trusting the story and start reading the log.

The tools compose into a sequence: understand β†’ verify β†’ measure β†’ control β†’ judge.

devin-internals-spec β€” understand. Before you can trust tool output you need to know the contracts: the file formats, exit codes, and behavioral rules of the runtime itself. This is the spec layer β€” everything downstream reads against it.

devin-qa-pack β€” verify. Runs a QA audit over actual agent work: file diffs present, tests run, commits exist, pushes landed, verification commands executed. Claims are checked against tool_call_state, not against the agent's narrative. 47 tests, CI on Ubuntu + Windows. This is the tool that started it all.

devin-evals β€” measure. Once verification works, you can ask the harder question: how good is the agent on this kind of task? Golden tasks, rubric scoring, regression tracking across sessions.

devin-bridge β€” control. A policy gate between intent and execution β€” ACP-based control with a --devin-only mode that enforces "this session does exactly what it was scoped to do."

poordjaevin β€” judge. The piece I'm proudest of. Takes a task description and produces a calibrated confidence score: should I trust this delegation? The calibration story is the interesting part β€” the judge went from ECE 0.170 to 0.071 through iterative refinement on real session data. (ECE = expected calibration error: when the judge says "80% confident," it should be right ~80% of the time. Most confidence scores don't do this. Now mine roughly does.) It's also a real MCP server β€” poordjaevin serve β€” so agents can consult it mid-flight.

Every tool follows the same contract: read local agent state, write local artifacts, expose a stable CLI, degrade gracefully when optional infrastructure is absent. No required daemons, no SaaS dashboard, no telemetry.

Concretely:

devin-history β€” cross-session memory: searchable SQLite over all past sessions (7+ commands of grep-able agent archaeology)devin-metrics β€” telemetry aggregation for quality signals over timedevin-memory + devin-search + devin-graph β€” memory store, retrieval, and a knowledge graph over sessions, projects and decisionsdevin-doctor β€” environment diagnostics across Windows/Linuxdevin-backup β€” snapshot + verify + restore of agent statedevin-janitor + devin-redact β€” cleanup and secret-redaction so state can move safelydevin-office β€” the fun one: a live circuit-board dashboard that renders real sessions, subagents and tool calls from the local store. Purely visual, read-only, zero telemetry β€” and it makes for a great demo GIF. 1. The agent's own telemetry is an untapped QA datasource. Session files and tool-call state are structured evidence most people ignore. Reading them turns "trust the agent" into "verify the agent" β€” and verification is what makes delegation safe at scale.

2. Calibration > confidence. A judge that's confidently wrong is worse than no judge. Iterating on ECE (0.170 β†’ 0.071) took real labeled outcomes, not prompt tweaks. If you build any kind of AI decision layer, measure calibration explicitly.

3. Devin-only mode matters more than integrations. Every tool works standalone with just the agent's CLI present. Obsidian, Slack, MCP β€” all optional. The ecosystem must survive a locked-down corporate box with nothing but Devin installed, because that's where a lot of real work happens.

Everything is MIT-licensed, cross-platform, and documented in English + PT-BR: devin-qa-pack (the verifier) or poordjaevin (the calibrated judge β€” pip install poordjaevin / uv tool install poordjaevin) Honest question for the comments: if you run AI agents on real work β€” Copilot, Devin, Cursor, Claude Code β€” how do you verify their claims today? Manual spot-checks? CI gates? Nothing? I suspect "nothing" is the most common answer, and I built this stack partly because it scared me that it was mine.

── more in #ai-agents 4 stories Β· sorted by recency
── more on @devin 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/i-built-tools-to-ver…] indexed:0 read:4min 2026-10-06 Β· β€”