{"slug": "two-prompts-for-one-coding-agent-task-implementation-and-supervision", "title": "Two prompts for one coding-agent task: implementation and supervision", "summary": "A developer outlined a dual-prompt pattern for coding-agent tasks that splits a single long prompt into a precise implementation brief and a shorter supervision brief, separating what the agent executes from who judges acceptance. The implementation brief defines scope, verification commands, and explicit pause points, while the supervision brief assigns review gates for scope, evidence, and domain expertise, with release decisions reserved for a human owner. The author notes the illustrative example does not prove two prompts prevent every mistake, but makes stop conditions visible before the run.", "body_md": "*AI assistance disclosure: This article was prepared with AI assistance. The example is illustrative.*\n\nA coding task often has one long prompt that mixes implementation instructions with managerial decisions. That makes the handoff brittle. The implementer needs a precise scope and tests; the reviewer needs a short record of what changed, what evidence exists, and when to stop the run. These are different documents for different readers.\n\nHere is a small example: a team wants a feature flag around a new checkout summary. The flag must default off. A coding agent can prepare the change, but the release owner must decide whether it goes live. This is an illustrative task template, not a report of a completed rollout.\n\n```\nTask: Put the new checkout summary behind `checkout_summary_v2`.\nScope: Change only the summary component, its flag lookup, and focused tests.\nStarting point: Work from a clean branch. First locate the current summary\nand the project's flag abstraction; report both paths before editing.\nBehavior: With the flag off, keep the current summary. With it on, show the\nnew summary. Keep totals, tax, and accessibility labels unchanged.\nVerification: Run the focused component tests and the relevant type check.\nRecord exact commands, exit codes, and any test failures. If you cannot run\nthem, say why; do not infer success from code inspection.\nPause when: the existing flag abstraction is absent; totals calculation\nwould change; tests require unavailable credentials; or a dependency needs\nan upgrade outside scope. Ask the task owner for a decision.\nDeliver: diff, test evidence, remaining risks, and a rollback note.\nDo not merge, enable the flag, or deploy.\n```\n\nThe brief supplies the agent with observable boundaries. A pause point is an instruction to stop and ask, not an invitation to choose a convenient workaround. The runner uses its own authorized repository and tool access. The person preparing this brief does not provide their login or tokens to the runner.\n\n```\nOwner: product engineer who prepared the task.\nReview gate 1 — scope: Does the diff touch only the approved files? If not,\nrequest a narrower patch or an explicit scope decision.\nReview gate 2 — evidence: Are the test commands and outputs attached? Re-run\nthe critical checkout path in the review environment.\nReview gate 3 — expertise: Have a payments-qualified reviewer examine any\nchange to price, tax, or transaction semantics. The task owner alone should\nnot approve those changes.\nAcceptance: Reviewer records pass/fail against the flag-off baseline and\nflag-on behavior. A delivered patch is not an accepted release.\nRelease: Only the release owner merges, schedules enablement, observes\nmetrics, and decides whether to roll back.\n```\n\nThis note is intentionally shorter. It says who judges acceptance and names decisions that an agent's test output cannot settle. A product or engineering lead prepares the handoff; the agent or teammate executes with their own permissions; a qualified human reviews the evidence and signs off. The separation is the core of human supervised AI execution.\n\nSuppose the runner finds that `checkout_summary_v2` already exists. It changes the component and adds two tests: flag off retains the old layout; flag on renders the new layout. It then discovers a screenshot baseline differs because a test fixture has an obsolete tax label. The correct next action is a pause with the failed command and diff attached. Updating tax labels quietly would cross the brief's boundary.\n\nThe reviewer can now distinguish three outcomes: implementation is ready for review, evidence is incomplete, or the task needs a new decision. None is the same as deployment. That is why [this dual-prompt design for human and agent tasks](https://wagglet.com/blog/dual-prompt-human-agent-task-design) is useful context for teams building a handoff process.\n\nThe example does not prove that two prompts prevent every mistake. It does make the stop condition visible before the work begins. For a small change, the two notes may fit in one task. For a risky change, add a named reviewer, a test environment, and a rollback owner before anyone claims it.", "url": "https://wpnews.pro/news/two-prompts-for-one-coding-agent-task-implementation-and-supervision", "canonical_source": "https://dev.to/wagglet/two-prompts-for-one-coding-agent-task-implementation-and-supervision-19j9", "published_at": "2026-10-08 06:37:42+00:00", "updated_at": "2026-10-08 06:46:50.698725+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "ai-safety"], "entities": [], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/two-prompts-for-one-coding-agent-task-implementation-and-supervision", "markdown": "https://wpnews.pro/news/two-prompts-for-one-coding-agent-task-implementation-and-supervision.md", "text": "https://wpnews.pro/news/two-prompts-for-one-coding-agent-task-implementation-and-supervision.txt", "jsonld": "https://wpnews.pro/news/two-prompts-for-one-coding-agent-task-implementation-and-supervision.jsonld"}}