# A screenshot diff can't accept a UI change for you: a task template with visual reference, test plan and reviewer verdict

> Source: <https://dev.to/sam_novak_574b07811e18495/a-screenshot-diff-cant-accept-a-ui-change-for-you-a-task-template-with-visual-reference-test-37k6>
> Published: 2026-10-08 09:12:56+00:00

Teams handing UI changes to coding agents often add a screenshot diff to the acceptance criteria. The thinking goes: the agent can't judge whether the result looks right, so let pixels decide. That's half right. A diff is a good test, but it's a bad verdict.

This tutorial builds a small, real example to show where the line is. It has a task template with a visual reference, a test plan that includes an automated screenshot diff, and a reviewer verdict. Then it runs a worked handoff with two implementations of the same change. The component and the task are invented. The diff script and every number in it are real, run with headless Chrome and pixelmatch 7.1.0.

*Drafted with AI assistance and reviewed by hand.*

A "Draft saved" toast at the bottom of the screen. The designer's reference covers three states (a short message, a long message, and a message with an Undo action) at two widths, 375 px and 1280 px. The long message must stop at two lines with an ellipsis.

The preparer writes this before anyone runs an agent:

```
T-512  "Draft saved" toast: new layout

Visual reference
  reference/toast.html, approved by: design (Mira), 2026-10-06
  States: default, long, undo      Widths: 375, 1280
  Exact:  colours, radius, padding, Undo placement
  Rule:   message clamps at 2 lines, ellipsis, at every width

Test plan
  Automated (runner attaches output)
    - unit: toast renders each state
    - visual: node check.mjs <build>, 3 states x 2 widths,
      fail above 0.5% changed pixels
  Manual (reviewer)
    - screen reader announces the message once (role="status")
    - Undo is reachable by keyboard and has a visible focus ring

Delivery evidence
  Branch, check.mjs output, screenshots of all 6 cells,
  anything the runner could not check and why

Reviewer verdict  (one line per row; pass / fail / not checked)
  [ ] visual: default    [ ] visual: long    [ ] visual: undo
  [ ] screen reader      [ ] keyboard
  Decision: accept | send back (with reason) | accept with noted deviation
```

Three people touch this, and they don't swap jobs:

That split is most of what project management for coding agents means in practice. The agent can produce a delivery. A person still decides whether it's accepted.

`check.mjs` takes a screenshot of every state at every width, for both the reference and the build, and compares them with pixelmatch. The core loop:

``` js
for (const state of STATES) {          // default, long, undo
  for (const width of WIDTHS) {        // 375, 1280
    const ref = shoot("reference", state, width);
    const got = shoot(impl, state, width);
    const changed = pixelmatch(ref.data, got.data, diff.data,
      ref.width, ref.height, { threshold: 0.1 });
    const share = changed / (ref.width * ref.height);
    console.log(`${share <= 0.005 ? "pass" : "FAIL"}  ${state} ${width}px ` +
      `${changed} px changed (${(share * 100).toFixed(2)}%)`);
  }
}
```

One trap I hit while building it: headless Chrome won't lay out a window as narrow as 375 px. My first run produced "375 px" screenshots that were really cut-off wider layouts. The fix was to render each state inside a frame of the exact width in a wider window, then crop to the frame. If your mobile screenshots look suspiciously like desktop ones, check this first.

The reference compared with itself gives 0 changed pixels in all six cells, so the check is stable on one machine.

Two deliveries for T-512. I wrote both by hand to stand in for agent output, and each is one CSS line away from the reference.

**Implementation A** drops the two-line clamp on the message. **Implementation B** keeps the clamp but uses 14 px vertical padding instead of 12 px.

The check's output:

```
--- impl-a
pass  default  375px  0 px changed (0.00%)
pass  default 1280px  0 px changed (0.00%)
FAIL  long     375px  15665 px changed (20.89%)
FAIL  long    1280px  13236 px changed (5.17%)
pass  undo     375px  0 px changed (0.00%)
pass  undo    1280px  0 px changed (0.00%)
--- impl-b
FAIL  default  375px  1667 px changed (2.22%)
FAIL  default 1280px  2103 px changed (0.82%)
FAIL  long     375px  3363 px changed (4.48%)
FAIL  long    1280px  4686 px changed (1.83%)
FAIL  undo     375px  1799 px changed (2.40%)
FAIL  undo    1280px  2235 px changed (0.87%)
```

Read it as a pixel count and B looks worse: 6 failures against A's 2. Read it as a reviewer and it's the other way round:

So the reviewer writes:

```
T-512 verdict, impl-a
  visual: default pass   visual: long FAIL   visual: undo pass
  screen reader: not checked   keyboard: not checked
  Decision: send back. Long message must clamp at 2 lines (rule in
  reference). Diff output attached.

T-512 verdict, impl-b
  visual: default/long/undo FAIL in check.mjs, 0.8% to 4.5%
  Cause: 14 px vertical padding vs 12 px in reference
  screen reader: ______   keyboard: ______   (reviewer runs these)
  Decision: if both manual rows pass, accept with noted deviation:
  14 px is the shared spacing token, so design updates the reference.
```

The manual rows for B are left blank on purpose. I didn't run a screen reader or a keyboard pass for this post, and a verdict template shouldn't arrive pre-filled. Two things here are the reason the verdict exists. "Not checked" is a valid entry: A was sent back before anyone checked the screen reader, and the verdict says so instead of leaving it blank. And "accept with noted deviation" records a human decision to overrule a failing check, with a reason someone can disagree with later.

Wagglet's [request-to-delivery workflow guide](https://wagglet.com/blog/wagglet-workflow-request-draft-ticket-delivery) treats a delivery and its acceptance as separate events with separate owners, which is the same split this template writes down for one UI task.
