cd /news/ai-agents/don-t-verify-the-agent-verify-the-st… Β· home β€Ί topics β€Ί ai-agents β€Ί article
[ARTICLE Β· art-146547] src=dev.to β†— pub= topic=ai-agents verified=true sentiment=Β· neutral

Don't verify the agent. Verify the state.

A developer building StareBrain, a pre-launch Android AI agent that requires user confirmation before executing actions, argues that verifying an agent's own success reports creates a recursive trust problem and instead proposes verifying system state directly. The approach checks ground truth outside the agent β€” such as whether a sent message exists in the sent folder with a timestamp after dispatch β€” and introduces DENIED_UNRESOLVED as a permanent first-class status for actions whose effects cannot be observed, rather than defaulting to a false success signal.

by read2 min views3 publishedOct 7, 2026

Building an AI agent that confirms before it acts forced us to confront a problem we didn't expect: how do you verify that an action actually happened?

The obvious answer is: check the agent's output. Ask it to confirm what it did.

That's wrong.

The recursive trust problem

If an agent reports "done" and you verify that report with another agent, you've just added a layer without solving anything. The second agent can be wrong for the same reasons the first one was. You haven't broken the trust chain β€” you've extended it.

Action dispatched
  β†’ Agent reports: "done"
    β†’ Verification agent checks: "looks done"
      β†’ System reports: success βœ“

Every step trusts the previous step's output. None of them look at the world.

What actually works: state diff

Instead of asking "what did the agent do," ask "did the system change the way we expected?"

Action dispatched: "send SMS to Sarah"
  β†’ Check sent folder: message present? βœ“
  β†’ Compare timestamp: after dispatch? βœ“
  β†’ System reports: confirmed βœ“

The verification is independent of the agent. You're not asking the agent to grade its own work β€” you're reading a ground truth that exists outside the agent entirely.

When state diff isn't possible

Some actions don't leave an observable state change. For those, the honest answer isn't "success" or "failure." It's UNRESOLVED.

We built DENIED_UNRESOLVED as a permanent first-class status in StareBrain β€” not a temporary placeholder that decays into an answer, but an explicit signal that says: the action was dispatched, but we cannot confirm what happened.

The confirmation screen surfaces this to the user:

"You'll know if this worked" β€” state is observable

"You might not know, and here's why" β€” state is not observable

Why this matters for AI agents specifically

An AI agent that confidently reports success on an unverifiable action is worse than one that reports nothing. Silence signals uncertainty. False confidence removes that signal entirely.

The verification layer has to be outside the trust chain. State diff gets you there for most actions. UNRESOLVED handles the rest honestly.

StareBrain is an Android AI agent: say what you want done, see exactly what it's about to do, confirm before anything runs. Pre-launch β€” waitlist open.

── more in #ai-agents 4 stories Β· sorted by recency
── more on @starebrain 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/don-t-verify-the-age…] indexed:0 read:2min 2026-10-07 Β· β€”