AI assistance disclosure: This article was prepared with AI assistance. The example is illustrative.
A coding task often has one long prompt that mixes implementation instructions with managerial decisions. That makes the handoff brittle. The implementer needs a precise scope and tests; the reviewer needs a short record of what changed, what evidence exists, and when to stop the run. These are different documents for different readers.
Here is a small example: a team wants a feature flag around a new checkout summary. The flag must default off. A coding agent can prepare the change, but the release owner must decide whether it goes live. This is an illustrative task template, not a report of a completed rollout.
Task: Put the new checkout summary behind `checkout_summary_v2`.
Scope: Change only the summary component, its flag lookup, and focused tests.
Starting point: Work from a clean branch. First locate the current summary
and the project's flag abstraction; report both paths before editing.
Behavior: With the flag off, keep the current summary. With it on, show the
new summary. Keep totals, tax, and accessibility labels unchanged.
Verification: Run the focused component tests and the relevant type check.
Record exact commands, exit codes, and any test failures. If you cannot run
them, say why; do not infer success from code inspection.
when: the existing flag abstraction is absent; totals calculation
would change; tests require unavailable credentials; or a dependency needs
an upgrade outside scope. Ask the task owner for a decision.
Deliver: diff, test evidence, remaining risks, and a rollback note.
Do not merge, enable the flag, or deploy.
The brief supplies the agent with observable boundaries. A point is an instruction to stop and ask, not an invitation to choose a convenient workaround. The runner uses its own authorized repository and tool access. The person preparing this brief does not provide their login or tokens to the runner.
Owner: product engineer who prepared the task.
Review gate 1 β scope: Does the diff touch only the approved files? If not,
request a narrower patch or an explicit scope decision.
Review gate 2 β evidence: Are the test commands and outputs attached? Re-run
the critical checkout path in the review environment.
Review gate 3 β expertise: Have a payments-qualified reviewer examine any
change to price, tax, or transaction semantics. The task owner alone should
not approve those changes.
Acceptance: Reviewer records pass/fail against the flag-off baseline and
flag-on behavior. A delivered patch is not an accepted release.
Release: Only the release owner merges, schedules enablement, observes
metrics, and decides whether to roll back.
This note is intentionally shorter. It says who judges acceptance and names decisions that an agent's test output cannot settle. A product or engineering lead prepares the handoff; the agent or teammate executes with their own permissions; a qualified human reviews the evidence and signs off. The separation is the core of human supervised AI execution.
Suppose the runner finds that checkout_summary_v2 already exists. It changes the component and adds two tests: flag off retains the old layout; flag on renders the new layout. It then discovers a screenshot baseline differs because a test fixture has an obsolete tax label. The correct next action is a with the failed command and diff attached. Updating tax labels quietly would cross the brief's boundary.
The reviewer can now distinguish three outcomes: implementation is ready for review, evidence is incomplete, or the task needs a new decision. None is the same as deployment. That is why this dual-prompt design for human and agent tasks is useful context for teams building a handoff process.
The example does not prove that two prompts prevent every mistake. It does make the stop condition visible before the work begins. For a small change, the two notes may fit in one task. For a risky change, add a named reviewer, a test environment, and a rollback owner before anyone claims it.