Teams handing UI changes to coding agents often add a screenshot diff to the acceptance criteria. The thinking goes: the agent can't judge whether the result looks right, so let pixels decide. That's half right. A diff is a good test, but it's a bad verdict.
This tutorial builds a small, real example to show where the line is. It has a task template with a visual reference, a test plan that includes an automated screenshot diff, and a reviewer verdict. Then it runs a worked handoff with two implementations of the same change. The component and the task are invented. The diff script and every number in it are real, run with headless Chrome and pixelmatch 7.1.0.
Drafted with AI assistance and reviewed by hand.
A "Draft saved" toast at the bottom of the screen. The designer's reference covers three states (a short message, a long message, and a message with an Undo action) at two widths, 375 px and 1280 px. The long message must stop at two lines with an ellipsis.
The preparer writes this before anyone runs an agent:
T-512 "Draft saved" toast: new layout
Visual reference
reference/toast.html, approved by: design (Mira), 2026-10-06
States: default, long, undo Widths: 375, 1280
Exact: colours, radius, padding, Undo placement
Rule: message clamps at 2 lines, ellipsis, at every width
Test plan
Automated (runner attaches output)
- unit: toast renders each state
- visual: node check.mjs <build>, 3 states x 2 widths,
fail above 0.5% changed pixels
Manual (reviewer)
- screen reader announces the message once (role="status")
- Undo is reachable by keyboard and has a visible focus ring
Delivery evidence
Branch, check.mjs output, screenshots of all 6 cells,
anything the runner could not check and why
Reviewer verdict (one line per row; pass / fail / not checked)
[ ] visual: default [ ] visual: long [ ] visual: undo
[ ] screen reader [ ] keyboard
Decision: accept | send back (with reason) | accept with noted deviation
Three people touch this, and they don't swap jobs:
That split is most of what project management for coding agents means in practice. The agent can produce a delivery. A person still decides whether it's accepted.
check.mjs takes a screenshot of every state at every width, for both the reference and the build, and compares them with pixelmatch. The core loop:
for (const state of STATES) { // default, long, undo
for (const width of WIDTHS) { // 375, 1280
const ref = shoot("reference", state, width);
const got = shoot(impl, state, width);
const changed = pixelmatch(ref.data, got.data, diff.data,
ref.width, ref.height, { threshold: 0.1 });
const share = changed / (ref.width * ref.height);
console.log(`${share <= 0.005 ? "pass" : "FAIL"} ${state} ${width}px ` +
`${changed} px changed (${(share * 100).toFixed(2)}%)`);
}
}
One trap I hit while building it: headless Chrome won't lay out a window as narrow as 375 px. My first run produced "375 px" screenshots that were really cut-off wider layouts. The fix was to render each state inside a frame of the exact width in a wider window, then crop to the frame. If your mobile screenshots look suspiciously like desktop ones, check this first.
The reference compared with itself gives 0 changed pixels in all six cells, so the check is stable on one machine.
Two deliveries for T-512. I wrote both by hand to stand in for agent output, and each is one CSS line away from the reference.
Implementation A drops the two-line clamp on the message. Implementation B keeps the clamp but uses 14 px vertical padding instead of 12 px.
The check's output:
--- impl-a
pass default 375px 0 px changed (0.00%)
pass default 1280px 0 px changed (0.00%)
FAIL long 375px 15665 px changed (20.89%)
FAIL long 1280px 13236 px changed (5.17%)
pass undo 375px 0 px changed (0.00%)
pass undo 1280px 0 px changed (0.00%)
--- impl-b
FAIL default 375px 1667 px changed (2.22%)
FAIL default 1280px 2103 px changed (0.82%)
FAIL long 375px 3363 px changed (4.48%)
FAIL long 1280px 4686 px changed (1.83%)
FAIL undo 375px 1799 px changed (2.40%)
FAIL undo 1280px 2235 px changed (0.87%)
Read it as a pixel count and B looks worse: 6 failures against A's 2. Read it as a reviewer and it's the other way round:
So the reviewer writes:
T-512 verdict, impl-a
visual: default pass visual: long FAIL visual: undo pass
screen reader: not checked keyboard: not checked
Decision: send back. Long message must clamp at 2 lines (rule in
reference). Diff output attached.
T-512 verdict, impl-b
visual: default/long/undo FAIL in check.mjs, 0.8% to 4.5%
Cause: 14 px vertical padding vs 12 px in reference
screen reader: ______ keyboard: ______ (reviewer runs these)
Decision: if both manual rows pass, accept with noted deviation:
14 px is the shared spacing token, so design updates the reference.
The manual rows for B are left blank on purpose. I didn't run a screen reader or a keyboard pass for this post, and a verdict template shouldn't arrive pre-filled. Two things here are the reason the verdict exists. "Not checked" is a valid entry: A was sent back before anyone checked the screen reader, and the verdict says so instead of leaving it blank. And "accept with noted deviation" records a human decision to overrule a failing check, with a reason someone can disagree with later.
Wagglet's request-to-delivery workflow guide treats a delivery and its acceptance as separate events with separate owners, which is the same split this template writes down for one UI task.