A coding agent can run tests, read the failure, edit code, and run the tests again. That sounds like a simple loop. Without a control structure, though, the same loop can become a string of untracked edits, repeated commands, and increasingly confident summaries.
A dependable workflow treats every action as a hypothesis about the system and every result as evidence that may change the next step. It also names the points where an agent must stop. The goal is not to keep the loop running until it says “done”; the goal is to produce a useful state transition with evidence a reviewer can inspect.
An engineering episode can be represented as a small state machine:
| State | Question to answer | Exit condition |
|---|---|---|
| Intake | What outcome, scope, and constraints were requested? | The task contract is understood; material ambiguity is recorded |
| Observe | What do the repository, tests, logs, and docs actually show? | Relevant facts and unknowns are separated |
| Model | What failure or change mechanism best explains the evidence? | There is a testable working hypothesis |
| Plan | What bounded actions can test that hypothesis? | Actions fit the authority and scope |
| Act | What changed, and what commands or tools ran? | The planned action produced a result or a stop condition |
| Evaluate | Does the result support the acceptance conditions? | Evidence is accepted, contradicted, stale, or insufficient |
| Exit | Is the episode complete, blocked, or ready for human judgment? | A truthful status and evidence report are returned |
The names are not sacred. A team can combine or rename states. What matters is that observation is not confused with inference, an attempted action is not confused with success, and a model’s summary is not substituted for a tool result.
Suppose a test fails after a dependency update. A disciplined trace might say:
The separation is useful because a plausible explanation can be wrong. If the agent records its hypothesis as though the logs established it, later actions inherit a false premise.
When multiple explanations remain possible, design the smallest useful experiment. Read the relevant library documentation, isolate the test, compare the old and new behavior, or inspect a sanitized trace. Avoid changing several unrelated variables at once; otherwise the outcome may not tell you which change mattered.
Verification is tied to the thing that was checked. If the code changes after a test run, that test result no longer describes the current code. If a test command is rerun with a different configuration, the old and new results are not interchangeable. If an agent changes the acceptance test itself, a green result needs a separate review of that change.
Treat an evidence item as a record with at least:
| Evidence field | Example |
|---|---|
| Subject | Working-tree revision or built artifact |
| Check | Exact command, test suite, or policy |
| Context | Runtime, fixture, configuration, and relevant dependencies |
| Result | Exit status plus failures, skips, or unknowns |
| Time/order | Whether it ran before or after the last relevant change |
An important rule follows: after a material edit, rerun the checks that the edit could invalidate. A test from before the patch can explain the starting condition; it cannot qualify the final patch.
This does not require rerunning every expensive job after every keystroke. Match the check to the risk and the changed surface. A documentation-only change might need a link check and rendered preview. A serialization change may need compatibility fixtures and migration tests. The completion report should say what was not rerun and why.
Retries are useful when the next attempt differs in a way that could address the observed failure. Repeating an identical command against an unchanged state usually gives the same evidence and consumes attention.
Before a retry, ask:
If a command fails because the local service is not running, starting that service is a meaningful next step. If it fails with the same assertion after two unmodified runs, another identical run is not a plan. If a tool call may have partially changed an external system, first determine whether repeating it is safe; a non-idempotent action can create duplicate tickets, releases, or payments. Bound retries by both count and consequence. For example: retry a deterministic local check once after a relevant fix; stop after a repeated infrastructure failure; require a person before repeating any external action whose first outcome is unknown.
Stop and escalate when:
These are workflow outcomes, not model moods. “I am confident” should not override a failed check or missing permission. Likewise, low confidence alone need not block a reversible, low-risk investigation if the agent can gather better evidence within its scope.
A useful trace records the task identifier, relevant inputs, tool calls, changed paths, check results, handoffs, and final status. Keep secrets and unnecessary personal data out of that record. The purpose is not to preserve every token; it is to let someone answer what the agent saw, what it did, what it learned, and why it stopped.
For teams using an agent platform, inspect what its trace actually captures. OpenAI’s current evaluation guide describes traces as end-to-end records of model calls, tool calls, guardrails, and handoffs. That is a useful example of a trace surface, but no vendor trace by itself proves that the application’s acceptance criteria were valid or that a consequential action was authorized. At the end of the episode, return a short, structured report:
This format makes a failed episode useful too. “Blocked because the provider’s migration guide leaves token refresh behavior unspecified” is actionable. “Couldn’t finish” is not.
Before adopting a workflow, walk through one ordinary task and one deliberately awkward case. Ask whether it can distinguish a fact from a hypothesis, invalidate old evidence, limit repeated actions, and reach a truthful stop state. If its only terminal condition is a success message, the loop is incomplete.
An agentic engineering loop is effective when feedback changes what happens next and authority constrains what may happen at all. That is how repeated tool use becomes a controlled engineering process instead of a long conversation with a terminal.
This article is based on the agentic engineering loop and evidence concepts in Part II of From Vibe Coding to Agentic Software Engineering, whose source record credits ChatGPT as preparer and identifies CC BY-NC-SA 4.0. This version is substantially reorganized and expanded with original examples and a practical workflow, and is shared under the same license: CC BY-NC-SA 4.0.
AI disclosure: The article text was generated primarily by AI. A human publisher supplied the topic, source material, and editorial direction, and remains responsible for checking claims and examples before publication.