# Two prompts for one coding-agent task: implementation and supervision

> Source: <https://dev.to/wagglet/two-prompts-for-one-coding-agent-task-implementation-and-supervision-19j9>
> Published: 2026-10-08 06:37:42+00:00

*AI assistance disclosure: This article was prepared with AI assistance. The example is illustrative.*

A coding task often has one long prompt that mixes implementation instructions with managerial decisions. That makes the handoff brittle. The implementer needs a precise scope and tests; the reviewer needs a short record of what changed, what evidence exists, and when to stop the run. These are different documents for different readers.

Here is a small example: a team wants a feature flag around a new checkout summary. The flag must default off. A coding agent can prepare the change, but the release owner must decide whether it goes live. This is an illustrative task template, not a report of a completed rollout.

```
Task: Put the new checkout summary behind `checkout_summary_v2`.
Scope: Change only the summary component, its flag lookup, and focused tests.
Starting point: Work from a clean branch. First locate the current summary
and the project's flag abstraction; report both paths before editing.
Behavior: With the flag off, keep the current summary. With it on, show the
new summary. Keep totals, tax, and accessibility labels unchanged.
Verification: Run the focused component tests and the relevant type check.
Record exact commands, exit codes, and any test failures. If you cannot run
them, say why; do not infer success from code inspection.
Pause when: the existing flag abstraction is absent; totals calculation
would change; tests require unavailable credentials; or a dependency needs
an upgrade outside scope. Ask the task owner for a decision.
Deliver: diff, test evidence, remaining risks, and a rollback note.
Do not merge, enable the flag, or deploy.
```

The brief supplies the agent with observable boundaries. A pause point is an instruction to stop and ask, not an invitation to choose a convenient workaround. The runner uses its own authorized repository and tool access. The person preparing this brief does not provide their login or tokens to the runner.

```
Owner: product engineer who prepared the task.
Review gate 1 — scope: Does the diff touch only the approved files? If not,
request a narrower patch or an explicit scope decision.
Review gate 2 — evidence: Are the test commands and outputs attached? Re-run
the critical checkout path in the review environment.
Review gate 3 — expertise: Have a payments-qualified reviewer examine any
change to price, tax, or transaction semantics. The task owner alone should
not approve those changes.
Acceptance: Reviewer records pass/fail against the flag-off baseline and
flag-on behavior. A delivered patch is not an accepted release.
Release: Only the release owner merges, schedules enablement, observes
metrics, and decides whether to roll back.
```

This note is intentionally shorter. It says who judges acceptance and names decisions that an agent's test output cannot settle. A product or engineering lead prepares the handoff; the agent or teammate executes with their own permissions; a qualified human reviews the evidence and signs off. The separation is the core of human supervised AI execution.

Suppose the runner finds that `checkout_summary_v2` already exists. It changes the component and adds two tests: flag off retains the old layout; flag on renders the new layout. It then discovers a screenshot baseline differs because a test fixture has an obsolete tax label. The correct next action is a pause with the failed command and diff attached. Updating tax labels quietly would cross the brief's boundary.

The reviewer can now distinguish three outcomes: implementation is ready for review, evidence is incomplete, or the task needs a new decision. None is the same as deployment. That is why [this dual-prompt design for human and agent tasks](https://wagglet.com/blog/dual-prompt-human-agent-task-design) is useful context for teams building a handoff process.

The example does not prove that two prompts prevent every mistake. It does make the stop condition visible before the work begins. For a small change, the two notes may fit in one task. For a risky change, add a named reviewer, a test environment, and a rollback owner before anyone claims it.
