cd /news/ai-agents/observe-model-act-verify-a-control-l… · home › topics › ai-agents › article
[ARTICLE · art-143375] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Observe, Model, Act, Verify: A Control Loop for Coding Agents

A developer outlined a control-loop framework for coding agents that models each engineering episode as a state machine spanning intake, observation, modeling, planning, action, evaluation, and exit. The approach requires agents to treat every action as a hypothesis and every result as evidence, keeping observation distinct from inference and rerunning checks invalidated by later edits. The writeup stresses that verification is tied to the specific revision and configuration tested, so results from before a material change cannot qualify the final patch.

by read6 min views4 publishedOct 1, 2026

A coding agent can run tests, read the failure, edit code, and run the tests again. That sounds like a simple loop. Without a control structure, though, the same loop can become a string of untracked edits, repeated commands, and increasingly confident summaries.

A dependable workflow treats every action as a hypothesis about the system and every result as evidence that may change the next step. It also names the points where an agent must stop. The goal is not to keep the loop running until it says “done”; the goal is to produce a useful state transition with evidence a reviewer can inspect.

An engineering episode can be represented as a small state machine:

State Question to answer Exit condition
Intake What outcome, scope, and constraints were requested? The task contract is understood; material ambiguity is recorded
Observe What do the repository, tests, logs, and docs actually show? Relevant facts and unknowns are separated
Model What failure or change mechanism best explains the evidence? There is a testable working hypothesis
Plan What bounded actions can test that hypothesis? Actions fit the authority and scope
Act What changed, and what commands or tools ran? The planned action produced a result or a stop condition
Evaluate Does the result support the acceptance conditions? Evidence is accepted, contradicted, stale, or insufficient
Exit Is the episode complete, blocked, or ready for human judgment? A truthful status and evidence report are returned

The names are not sacred. A team can combine or rename states. What matters is that observation is not confused with inference, an attempted action is not confused with success, and a model’s summary is not substituted for a tool result.

Suppose a test fails after a dependency update. A disciplined trace might say:

The separation is useful because a plausible explanation can be wrong. If the agent records its hypothesis as though the logs established it, later actions inherit a false premise.

When multiple explanations remain possible, design the smallest useful experiment. Read the relevant library documentation, isolate the test, compare the old and new behavior, or inspect a sanitized trace. Avoid changing several unrelated variables at once; otherwise the outcome may not tell you which change mattered.

Verification is tied to the thing that was checked. If the code changes after a test run, that test result no longer describes the current code. If a test command is rerun with a different configuration, the old and new results are not interchangeable. If an agent changes the acceptance test itself, a green result needs a separate review of that change.

Treat an evidence item as a record with at least:

Evidence field Example
Subject Working-tree revision or built artifact
Check Exact command, test suite, or policy
Context Runtime, fixture, configuration, and relevant dependencies
Result Exit status plus failures, skips, or unknowns
Time/order Whether it ran before or after the last relevant change

An important rule follows: after a material edit, rerun the checks that the edit could invalidate. A test from before the patch can explain the starting condition; it cannot qualify the final patch.

This does not require rerunning every expensive job after every keystroke. Match the check to the risk and the changed surface. A documentation-only change might need a link check and rendered preview. A serialization change may need compatibility fixtures and migration tests. The completion report should say what was not rerun and why.

Retries are useful when the next attempt differs in a way that could address the observed failure. Repeating an identical command against an unchanged state usually gives the same evidence and consumes attention.

Before a retry, ask:

If a command fails because the local service is not running, starting that service is a meaningful next step. If it fails with the same assertion after two unmodified runs, another identical run is not a plan. If a tool call may have partially changed an external system, first determine whether repeating it is safe; a non-idempotent action can create duplicate tickets, releases, or payments. Bound retries by both count and consequence. For example: retry a deterministic local check once after a relevant fix; stop after a repeated infrastructure failure; require a person before repeating any external action whose first outcome is unknown.

Stop and escalate when:

These are workflow outcomes, not model moods. “I am confident” should not override a failed check or missing permission. Likewise, low confidence alone need not block a reversible, low-risk investigation if the agent can gather better evidence within its scope.

A useful trace records the task identifier, relevant inputs, tool calls, changed paths, check results, handoffs, and final status. Keep secrets and unnecessary personal data out of that record. The purpose is not to preserve every token; it is to let someone answer what the agent saw, what it did, what it learned, and why it stopped.

For teams using an agent platform, inspect what its trace actually captures. OpenAI’s current evaluation guide describes traces as end-to-end records of model calls, tool calls, guardrails, and handoffs. That is a useful example of a trace surface, but no vendor trace by itself proves that the application’s acceptance criteria were valid or that a consequential action was authorized. At the end of the episode, return a short, structured report:

This format makes a failed episode useful too. “Blocked because the provider’s migration guide leaves token refresh behavior unspecified” is actionable. “Couldn’t finish” is not.

Before adopting a workflow, walk through one ordinary task and one deliberately awkward case. Ask whether it can distinguish a fact from a hypothesis, invalidate old evidence, limit repeated actions, and reach a truthful stop state. If its only terminal condition is a success message, the loop is incomplete.

An agentic engineering loop is effective when feedback changes what happens next and authority constrains what may happen at all. That is how repeated tool use becomes a controlled engineering process instead of a long conversation with a terminal.

This article is based on the agentic engineering loop and evidence concepts in Part II of From Vibe Coding to Agentic Software Engineering, whose source record credits ChatGPT as preparer and identifies CC BY-NC-SA 4.0. This version is substantially reorganized and expanded with original examples and a practical workflow, and is shared under the same license: CC BY-NC-SA 4.0.

AI disclosure: The article text was generated primarily by AI. A human publisher supplied the topic, source material, and editorial direction, and remains responsible for checking claims and examples before publication.

── more in #ai-agents 4 stories · sorted by recency
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/observe-model-act-ve…] indexed:0 read:6min 2026-10-01 · —