AI coding agents are great at one thing that isn't writing code: asserting. "Tests pass." "The file was updated." "I pushed the fix." And if you've run agents for real work, you know these claims are sometimes... optimistic.
I'm a QA engineer who runs Devin daily. At some point I got tired of manually checking whether the agent's claims matched reality, so I did what a QA engineer does: I built a test harness. Then an eval suite. Then a policy layer. Then a decision layer. Twenty local-first tools later, the whole stack runs on a RAM-constrained laptop with zero telemetry leaving the machine.
This is the short version of what exists, why, and what I learned.
The trigger was simple. An agent session reported "fixed, all green" β and the file on disk disagreed. No malice; agents report from their own narrative, not from ground truth. Classic QA problem: the claim and the state are different artifacts, and only one of them is evidence.
The insight that made everything else possible: agent CLIs already write structured telemetry locally β session files, tool-call state, token usage. The evidence was sitting on disk the whole time. I just had to stop trusting the story and start reading the log.
The tools compose into a sequence: understand β verify β measure β control β judge.
devin-internals-spec β understand. Before you can trust tool output you need to know the contracts: the file formats, exit codes, and behavioral rules of the runtime itself. This is the spec layer β everything downstream reads against it.
devin-qa-pack β verify. Runs a QA audit over actual agent work: file diffs present, tests run, commits exist, pushes landed, verification commands executed. Claims are checked against tool_call_state, not against the agent's narrative. 47 tests, CI on Ubuntu + Windows. This is the tool that started it all.
devin-evals β measure. Once verification works, you can ask the harder question: how good is the agent on this kind of task? Golden tasks, rubric scoring, regression tracking across sessions.
devin-bridge β control. A policy gate between intent and execution β ACP-based control with a --devin-only mode that enforces "this session does exactly what it was scoped to do."
poordjaevin β judge. The piece I'm proudest of. Takes a task description and produces a calibrated confidence score: should I trust this delegation? The calibration story is the interesting part β the judge went from ECE 0.170 to 0.071 through iterative refinement on real session data. (ECE = expected calibration error: when the judge says "80% confident," it should be right ~80% of the time. Most confidence scores don't do this. Now mine roughly does.) It's also a real MCP server β poordjaevin serve β so agents can consult it mid-flight.
Every tool follows the same contract: read local agent state, write local artifacts, expose a stable CLI, degrade gracefully when optional infrastructure is absent. No required daemons, no SaaS dashboard, no telemetry.
Concretely:
devin-history β cross-session memory: searchable SQLite over all past sessions (7+ commands of grep-able agent archaeology)devin-metrics β telemetry aggregation for quality signals over timedevin-memory + devin-search + devin-graph β memory store, retrieval, and a knowledge graph over sessions, projects and decisionsdevin-doctor β environment diagnostics across Windows/Linuxdevin-backup β snapshot + verify + restore of agent statedevin-janitor + devin-redact β cleanup and secret-redaction so state can move safelydevin-office β the fun one: a live circuit-board dashboard that renders real sessions, subagents and tool calls from the local store. Purely visual, read-only, zero telemetry β and it makes for a great demo GIF.
1. The agent's own telemetry is an untapped QA datasource. Session files and tool-call state are structured evidence most people ignore. Reading them turns "trust the agent" into "verify the agent" β and verification is what makes delegation safe at scale.
2. Calibration > confidence. A judge that's confidently wrong is worse than no judge. Iterating on ECE (0.170 β 0.071) took real labeled outcomes, not prompt tweaks. If you build any kind of AI decision layer, measure calibration explicitly.
3. Devin-only mode matters more than integrations. Every tool works standalone with just the agent's CLI present. Obsidian, Slack, MCP β all optional. The ecosystem must survive a locked-down corporate box with nothing but Devin installed, because that's where a lot of real work happens.
Everything is MIT-licensed, cross-platform, and documented in English + PT-BR:
devin-qa-pack (the verifier) or poordjaevin (the calibrated judge β pip install poordjaevin / uv tool install poordjaevin)
Honest question for the comments: if you run AI agents on real work β Copilot, Devin, Cursor, Claude Code β how do you verify their claims today? Manual spot-checks? CI gates? Nothing? I suspect "nothing" is the most common answer, and I built this stack partly because it scared me that it was mine.