Your agent loop is not a production system AWS DevOps Agent's directed actions, which create, modify, or otherwise mutate resources, are disabled by default and require layered opt-in and per-action approval, executing with credentials scoped to the approved operation and resource and attributable to the approving operator in AWS CloudTrail, according to an account from an engineer who worked on the system. The engineer argues that approval, execution, and completion are three different proofs, since a tool can run successfully without the operational problem being solved, and that leaders should standardize five contracts — context, state, tools, control, and completion — before teams scale agent autonomy. Mutation, not model choice, is the production boundary: once an agent can change state, approval, identity, attribution, and verification must become first-class system concerns. You have 1 article left to read this month before you need to register https://leaddev.com/register a free LeadDev.com account. Estimated reading time: 12 minutes Key takeaways - Mutation, not model choice, is the production boundary . Once an agent can change state, approval, identity, attribution, and verification have to become first-class system concerns. - Approval, execution, and completion are three different proofs. A tool can run successfully without the operational problem being solved . - Standardize the harness contracts centrally , while domain teams retain ownership of tool semantics, task policy, and the definition of done. Agent https://leaddev.com/technical-direction/how-to-prepare-for-ai-agents frameworks make it easy to connect models to tools. The harder work begins when those tools can change production. Drawing on my work on AWS DevOps Agent, I argue that leaders should standardize five contracts – context, state, tools, control, and completion – before teams scale agent autonomy. At 2 am, an incident response https://leaddev.com/technical-direction/incident-response-before-the-incident agent is not having a conversation. It is joining an on-call engineer in the middle of a production event. It may inspect an alert, correlate metrics and logs, compare recent deployments, pull a runbook, and test a few hypotheses. As long as the agent stays read-only, most bad ideas remain recommendations that an engineer can challenge. The architecture changes the moment the operator says, “apply it.” A wrong recommendation is recoverable. A wrong mutation has a blast radius. Your inbox, upgraded. Receive weekly engineering insights to level up your leadership approach. AWS DevOps Agent exposes this boundary through directed actions https://docs.aws.amazon.com/devopsagent/latest/userguide/working-with-devops-agent-working-with-directed-actions.html : operations that create, modify, or otherwise mutate resources. They are disabled by default, require layered opt-in and per-action approval, execute with credentials scoped to the approved operation and resource, and are attributable to the approving operator in AWS CloudTrail. Working on AWS DevOps Agent changed where I drew the system boundary. Approval could not be a transient answer attached to one conversation turn. Runs pause and resume. Targets can change. The same capability may be reachable through more than one execution path. The approval has to become a durable state with an exact scope, an identity, a lifetime, and an explicit reuse rule. Just as importantly, a successful tool response could prove that an action ran without proving that the incident was over. That was the point where the harness stopped looking like glue around a model and started looking like a control plane. I have been working through this problem from the bottom up. In Keep the Terminal Relevant: Patterns for AI Agent Driven CLIs https://www.infoq.com/articles/ai-agent-cli/ , I focused on the interfaces that let agents use software reliably. In Trustworthy Productivity: Securing AI Accelerated Development https://www.infoq.com/articles/secure-ai-development/ , I moved up to the ReAct loop and treated context, reasoning, and tools as separate trust boundaries. This is the next layer. Agent-ready tools answer whether an agent can use software reliably. Guardrails https://leaddev.com/software-quality/how-decide-engineering-guardrails answer whether it can act inside a defensible envelope. The harness answers whether those properties survive a real production run. The loop is only control flow Most agent demos reduce to a small loop: Model receives context - model proposes a tool call - tool returns a result - model receives another turn That loop is useful. It is not an operating model. A production harness has to preserve the chain around the model: which evidence shaped the proposal, what action the operator approved, who approved it, which identity executed it, what changed, and why the system continued or stopped. For engineering leaders https://leaddev.com/the-engineering-leadership-report-2026/ , this changes the question. The first decision is not whether to use Strands, LangGraph, or another toolkit. It is what guarantees every team should inherit before any agent receives write access. Standardize five contracts before teams scale I use five contracts: context https://leaddev.com/ai/what-is-context-engineering , state, tools, control, and completion. A framework component tells you that a checkpoint store or approval hook exists. A contract tells you what must remain true after a restart, a cancellation, a scope change, or a false claim of success. These contracts also create a clean ownership model. A platform team can own the durable action record, context interfaces, execution gateway, identity propagation, interruption, common observability, and evaluation substrate. Domain teams should own their evidence sources, tool semantics, task policies, and completion criteria. Security teams should own hard invariants and adversarial tests. Operators should own approval thresholds, escalation rules, and the ability to inspect and stop work. Centralize the contracts, not the intelligence. That gives teams a paved path without turning the harness into a monolith. Models and orchestration styles can change. The behavior that matters under stress should not. A successful action is not a successful task Consider a common incident pattern. An agent https://leaddev.com/technical-direction/why-everyones-suddenly-talking-about-ai-agents sees a sharp increase in HTTP 5xx errors shortly after a deployment. It correlates the timing with a configuration change, rules out an upstream dependency, and proposes rolling the checkout service back to the previous release. The operator should not be asked to approve the sentence “roll it back.” That is too vague. The approval surface should show the exact operation, target resource, material parameters, expected blast radius, and validity window. The operator can narrow the request, reject it, or approve it for a bounded use. More like this Once approved, the harness executes with scoped authority and records attribution. It then reads the deployment state back. Is the intended release actually active on the intended service? Even that is not enough. The harness still has to check the incident-level postcondition. Did the 5xx rate recover? Did the alarm clear? Does the proposed root cause still fit the evidence? Approval proves authority. Execution evidence proves what changed. Verification proves whether the goal was reached. These are easy to collapse in a demo. A confirmation box can look like authorization. A successful API response can look like proof. A completed rollback can look like a resolved incident. In production, they are three different facts. The five contracts are how the harness preserves those facts across the run. 1. Context: pin the decision Context is often assembled by concatenating system instructions, recent messages, retrieved documents, and tool results into one large prompt. That works until information comes from different trust boundaries or the world changes while the run is in progress. A policy https://leaddev.com/software-quality/every-decision-creates-policy is not the same thing as a runbook. A runbook is not the same thing as a log line. A log line that says “ignore the previous instruction” is still evidence, not an instruction. Each item needs provenance, freshness, task scope, and a clear semantic role. Mutation adds another requirement: the thing being approved has to stay stable. Approval for “roll back checkout” cannot silently attach itself to whichever release happens to be latest five minutes later. Record the exact tool https://leaddev.com/ai/best-ai-coding-assistants , operation, target, material parameters, evidence references, risk, requested identity, approval window, and expected postconditions. If an approval-relevant field changes, ask again. Failure test: change the target deployment after the approval request is created. The harness should use the pinned target or require fresh approval. It should never reinterpret the old decision against new context. 2. State: resume authority, not just chat Conversation history, runtime state, and model context are related. They are not interchangeable. Runtime state includes approvals, budgets https://leaddev.com/career-development/rethinking-your-engineering-budget-during-downturn , identities, checkpoints, execution records, verification results, and the current task graph. Model context is the smaller view selected for one inference. Put everything into a mutable message array and the system appears durable only while one process stays alive. The durable record does not need every model token. It needs the facts that govern the next legal step: what evidence was admitted, what was proposed, which policy allowed it, who approved it, which scope was finalized, whether that approval had already been used, what executed, and what verification established. If a process restarts after approval but before execution, the model should not reconstruct authority from chat. The harness should restore the exact scope and either resume the permitted action or ask again if the approval is no longer valid. Failure test: stop the client after approval is recorded but before execution. Resume on another host and prove that the scope is restored, used at most once, and never widened by a fresh model turn. If a harness can restore the conversation but not the authorization, it has memory, not durable state. 3. Tools: give every mutation one path A tool schema is only the starting point. For every tool, a production team eventually has to answer: is it read-only, mutating, or destructive? Which identity and resources does it need? Can it dry-run? How long can it run? What structured result comes back? How can the affected resource be read again? Which postcondition tells us the larger task succeeded? Classification has to match real behavior, and every invocation path has to respect it. A write reached through generated code is still a write. So is a write invoked through Model Context Protocol MCP https://leaddev.com/ai/an-engineers-guide-to-model-context-protocol-mcp , a background worker, or a child agent. Every mutation should pass through the same gateway: validate, classify, authorize, request approval, finalize scope, execute, read back, verify, and record. Once one path can skip that sequence, the harness has two security models and two reliability models. This is the lesson from agent-ready CLIs one layer up. Machine consumers need stable contracts, not prose that happens to look understandable. At the harness layer, the contract also includes identity, approval scope, structured execution evidence, resource read-back, and task-level proof. Failure test: invoke the same write through a direct tool, generated code, an MCP tool, and a child agent. Every path should receive the same classification, scoped authority, approval requirement, and audit record. 4. Control: make human decisions durable Human-in-the-loop https://leaddev.com/ai/designing-human-agent-engineering-teams is often implemented as a yes or no dialog bolted onto a tool call. In production, an approval is closer to a temporary capability. It should name the operation, resource, parameter constraints, approver, expiry, and reuse rules. A yes to update service A is not permission to update service B or repeat the action later. The final scope may be narrower than the proposal, never wider. The decision should be serializable so the run can pause and resume. It should be revocable until used and attributable after execution. Control also includes budgets for steps, time, cost, tool calls, and delegation depth. Cancellation should reach in-flight tools and child agents. Delegation should reduce authority, not expand it: a child agent needs a task-scoped tool set, bounded depth, explicit parentage, and a structured return channel. Failure test: cancel while a child agent is evaluating mitigations. Prove that cancellation propagates, no new mutation starts, and the durable state can be inspected before resuming. Human-in-the-loop is not a confirmation dialog. It is a durable authorization transition. 5. Completion: verify the operational outcome The model saying “done” is not a completion condition. A coding agent https://leaddev.com/ai/your-ai-coding-tools-buying-checklist-for-2026 may need passing tests, a clean diff, a successful build, and a runtime check. An incident-response agent may need the alarm to clear, key metrics to normalize, and the proposed root cause to explain the evidence. The definition of done belongs to the task, not the model. Where practical, separate the verifier from the planner. The planner is trying to find a path to the goal. The verifier asks whether the goal was reached and whether the evidence is good enough. Letting the same unchecked assertion do both creates a system that can confidently grade its own answer. A rollback can succeed while the incident continues. The deployment system may confirm that the previous version is active while the 5xx rate remains elevated. The directed action worked. The operational task did not. The harness should continue the investigation or escalate rather than declare victory because one API call completed. This also changes what leaders should measure. Tool-call count, token volume, and unattended run time are activity metrics. More useful measures include verified completion rate, correction and escalation rate, recovery after interruption, operator effort needed to understand a run, and cost per verified outcome. Failure test: make the model report that the incident is resolved while the alarm remains active. The harness should reject the terminal state and show the missing postcondition. Frameworks give you hooks, not an operating model Agent toolkits are converging on useful primitives. Strands https://strandsagents.com/docs/user-guide/concepts/interrupts/ has interrupts and interventions. LangGraph https://docs.langchain.com/oss/python/langgraph/interrupts checkpoints state and supports resumable interrupts. The OpenAI Agents SDK https://openai.github.io/openai-agents-python/human in the loop/ can serialize paused approval state. Google ADK https://google.github.io/adk-docs/runtime/event-loop/ exposes events, session state, and callbacks around agents, models, and tools. That is progress. Teams no longer have to invent every lifecycle hook, but a checkpoint does not define your approval semantics. A callback does not prove that alternate execution paths honor the same policy. A trace does not prove that the resource changed or that the task succeeded. Frameworks https://leaddev.com/management/level-frameworks-developing-developers provide interception points. Your organization still owns their meaning. Berlin • November 9 & 10, 2026 Close the gap between what leadership expects and what’s actually possible at LeadDev Berlin . Build one directed action before you scale autonomy The ecosystem makes it easy to start with plugins, generated code, background workers, and multi-agent collaboration. Those features make demos look powerful. They also multiply the places where state, authority, and completion can go wrong. I would build a new harness in three stages. 1. Make one read-only path reconstructable. Start with one model and one read-only tool. Add a durable event record, structured tool results, and a reliable stop mechanism. The first milestone is being able to explain the run, not making the agent look autonomous. 2. Add one operator-directed write. Require a concrete proposal, classification, task-scoped identity, per-action approval, resource read-back, and a verifiable postcondition. This is where the five contracts become real. 3. Break it before you expand it. Restart after approval, expire the decision, change the target, deny the action, force the tool to fail, cancel delegated work, and keep the postcondition false. Only after those paths are trustworthy should you add generated code, background work, plugins, or child agents. This order is deliberately boring. Boring is good. It forces the semantics to become clear before the number of execution paths explodes. Trustworthy productivity has to survive the action Models will change faster than the systems around them. Prompts will evolve. Tool protocols will expand. Today’s runtime will eventually be replaced. The durable asset is the operational judgment encoded in the harness: what enters context, what persists, how operator intent is pinned, how authority is narrowed, where humans can interrupt, how execution is attributed, and what counts as enough evidence to stop. This is the layer that connects my earlier arguments. Agent-ready interfaces make tools usable. Guardrails keep the agent inside a defensible envelope. The harness carries those properties through a real production run. Agent-ready tools make action possible. Guardrails make action bounded. The harness makes action inspectable, attributable, and verifiable. The question is not whether a model can call a tool. The question is whether the system can show what it was allowed to do, what it actually changed, and whether the result was good enough to stop. The loop makes the agent useful. The harness is what makes it fit for production.