# I built tools to verify what my AI coding agent actually did

> Source: <https://dev.to/icaro0310/i-built-tools-to-verify-what-my-ai-coding-agent-actually-did-2a6f>
> Published: 2026-10-06 17:12:11+00:00

AI coding agents are great at one thing that isn't writing code: *asserting*. "Tests pass." "The file was updated." "I pushed the fix." And if you've run agents for real work, you know these claims are sometimes... optimistic.

I'm a QA engineer who runs [Devin](https://devin.ai) daily. At some point I got tired of manually checking whether the agent's claims matched reality, so I did what a QA engineer does: I built a test harness. Then an eval suite. Then a policy layer. Then a decision layer. Twenty local-first tools later, the whole stack runs on a RAM-constrained laptop with zero telemetry leaving the machine.

This is the short version of what exists, why, and what I learned.

The trigger was simple. An agent session reported "fixed, all green" — and the file on disk disagreed. No malice; agents report from their own narrative, not from ground truth. Classic QA problem: the *claim* and the *state* are different artifacts, and only one of them is evidence.

The insight that made everything else possible: agent CLIs already write structured telemetry locally — session files, tool-call state, token usage. The evidence was sitting on disk the whole time. I just had to stop trusting the story and start reading the log.

The tools compose into a sequence: understand → verify → measure → control → judge.

**[devin-internals-spec](https://github.com/Icaro0310/devin-internals-spec)** — *understand.* Before you can trust tool output you need to know the contracts: the file formats, exit codes, and behavioral rules of the runtime itself. This is the spec layer — everything downstream reads against it.

**[devin-qa-pack](https://github.com/Icaro0310/devin-qa-pack)** — *verify.* Runs a QA audit over actual agent work: file diffs present, tests run, commits exist, pushes landed, verification commands executed. Claims are checked against `tool_call_state`, not against the agent's narrative. 47 tests, CI on Ubuntu + Windows. This is the tool that started it all.

**[devin-evals](https://github.com/Icaro0310/devin-evals)** — *measure.* Once verification works, you can ask the harder question: how good is the agent on *this* kind of task? Golden tasks, rubric scoring, regression tracking across sessions.

**[devin-bridge](https://github.com/Icaro0310/devin-bridge)** — *control.* A policy gate between intent and execution — ACP-based control with a `--devin-only` mode that enforces "this session does exactly what it was scoped to do."

**[poordjaevin](https://github.com/Icaro0310/poordjaevin)** — *judge.* The piece I'm proudest of. Takes a task description and produces a calibrated confidence score: should I trust this delegation? The calibration story is the interesting part — the judge went from **ECE 0.170 to 0.071** through iterative refinement on real session data. (ECE = expected calibration error: when the judge says "80% confident," it should be right ~80% of the time. Most confidence scores don't do this. Now mine roughly does.) It's also a real MCP server — `poordjaevin serve` — so agents can consult it mid-flight.

Every tool follows the same contract: read local agent state, write local artifacts, expose a stable CLI, degrade gracefully when optional infrastructure is absent. No required daemons, no SaaS dashboard, no telemetry.

Concretely:

`devin-history` — cross-session memory: searchable SQLite over all past sessions (7+ commands of grep-able agent archaeology)`devin-metrics` — telemetry aggregation for quality signals over time`devin-memory` + `devin-search` + `devin-graph` — memory store, retrieval, and a knowledge graph over sessions, projects and decisions`devin-doctor` — environment diagnostics across Windows/Linux`devin-backup` — snapshot + verify + restore of agent state`devin-janitor` + `devin-redact` — cleanup and secret-redaction so state can move safely`devin-office` — the fun one: a live circuit-board dashboard that renders real sessions, subagents and tool calls from the local store. Purely visual, read-only, zero telemetry — and it makes for a great demo GIF.
**1. The agent's own telemetry is an untapped QA datasource.** Session files and tool-call state are structured evidence most people ignore. Reading them turns "trust the agent" into "verify the agent" — and verification is what makes delegation safe at scale.

**2. Calibration > confidence.** A judge that's confidently wrong is worse than no judge. Iterating on ECE (0.170 → 0.071) took real labeled outcomes, not prompt tweaks. If you build any kind of AI decision layer, measure calibration explicitly.

**3. Devin-only mode matters more than integrations.** Every tool works standalone with just the agent's CLI present. Obsidian, Slack, MCP — all optional. The ecosystem must survive a locked-down corporate box with nothing but Devin installed, because that's where a lot of real work happens.

Everything is MIT-licensed, cross-platform, and documented in English + PT-BR:

`devin-qa-pack` (the verifier) or `poordjaevin` (the calibrated judge — `pip install poordjaevin` / `uv tool install poordjaevin`)
Honest question for the comments: if you run AI agents on real work — Copilot, Devin, Cursor, Claude Code — **how do you verify their claims today?** Manual spot-checks? CI gates? Nothing? I suspect "nothing" is the most common answer, and I built this stack partly because it scared me that it was mine.
