{"slug": "a-screenshot-diff-can-t-accept-a-ui-change-for-you-a-task-template-with-visual", "title": "A screenshot diff can't accept a UI change for you: a task template with visual reference, test plan and reviewer verdict", "summary": "A developer published a task-template workflow for handing UI changes to coding agents that pairs an automated screenshot diff with a human reviewer verdict, arguing that a pixel diff is a valid test but not a valid acceptance decision. The worked example runs headless Chrome and pixelmatch 7.1.0 across three toast states at 375 px and 1280 px with a 0.5% changed-pixel threshold, and shows two one-line CSS deviations producing different failure signatures — a dropped two-line clamp failing the long-message state at 20.89% and 5.17% changed pixels, versus a padding change caught elsewhere. The writeup also flags that headless Chrome will not lay out a 375 px window, so narrow screenshots must be rendered in a fixed-width frame and cropped.", "body_md": "Teams handing UI changes to coding agents often add a screenshot diff to the acceptance criteria. The thinking goes: the agent can't judge whether the result looks right, so let pixels decide. That's half right. A diff is a good test, but it's a bad verdict.\n\nThis tutorial builds a small, real example to show where the line is. It has a task template with a visual reference, a test plan that includes an automated screenshot diff, and a reviewer verdict. Then it runs a worked handoff with two implementations of the same change. The component and the task are invented. The diff script and every number in it are real, run with headless Chrome and pixelmatch 7.1.0.\n\n*Drafted with AI assistance and reviewed by hand.*\n\nA \"Draft saved\" toast at the bottom of the screen. The designer's reference covers three states (a short message, a long message, and a message with an Undo action) at two widths, 375 px and 1280 px. The long message must stop at two lines with an ellipsis.\n\nThe preparer writes this before anyone runs an agent:\n\n```\nT-512  \"Draft saved\" toast: new layout\n\nVisual reference\n  reference/toast.html, approved by: design (Mira), 2026-10-06\n  States: default, long, undo      Widths: 375, 1280\n  Exact:  colours, radius, padding, Undo placement\n  Rule:   message clamps at 2 lines, ellipsis, at every width\n\nTest plan\n  Automated (runner attaches output)\n    - unit: toast renders each state\n    - visual: node check.mjs <build>, 3 states x 2 widths,\n      fail above 0.5% changed pixels\n  Manual (reviewer)\n    - screen reader announces the message once (role=\"status\")\n    - Undo is reachable by keyboard and has a visible focus ring\n\nDelivery evidence\n  Branch, check.mjs output, screenshots of all 6 cells,\n  anything the runner could not check and why\n\nReviewer verdict  (one line per row; pass / fail / not checked)\n  [ ] visual: default    [ ] visual: long    [ ] visual: undo\n  [ ] screen reader      [ ] keyboard\n  Decision: accept | send back (with reason) | accept with noted deviation\n```\n\nThree people touch this, and they don't swap jobs:\n\nThat split is most of what project management for coding agents means in practice. The agent can produce a delivery. A person still decides whether it's accepted.\n\n`check.mjs` takes a screenshot of every state at every width, for both the reference and the build, and compares them with pixelmatch. The core loop:\n\n``` js\nfor (const state of STATES) {          // default, long, undo\n  for (const width of WIDTHS) {        // 375, 1280\n    const ref = shoot(\"reference\", state, width);\n    const got = shoot(impl, state, width);\n    const changed = pixelmatch(ref.data, got.data, diff.data,\n      ref.width, ref.height, { threshold: 0.1 });\n    const share = changed / (ref.width * ref.height);\n    console.log(`${share <= 0.005 ? \"pass\" : \"FAIL\"}  ${state} ${width}px ` +\n      `${changed} px changed (${(share * 100).toFixed(2)}%)`);\n  }\n}\n```\n\nOne trap I hit while building it: headless Chrome won't lay out a window as narrow as 375 px. My first run produced \"375 px\" screenshots that were really cut-off wider layouts. The fix was to render each state inside a frame of the exact width in a wider window, then crop to the frame. If your mobile screenshots look suspiciously like desktop ones, check this first.\n\nThe reference compared with itself gives 0 changed pixels in all six cells, so the check is stable on one machine.\n\nTwo deliveries for T-512. I wrote both by hand to stand in for agent output, and each is one CSS line away from the reference.\n\n**Implementation A** drops the two-line clamp on the message. **Implementation B** keeps the clamp but uses 14 px vertical padding instead of 12 px.\n\nThe check's output:\n\n```\n--- impl-a\npass  default  375px  0 px changed (0.00%)\npass  default 1280px  0 px changed (0.00%)\nFAIL  long     375px  15665 px changed (20.89%)\nFAIL  long    1280px  13236 px changed (5.17%)\npass  undo     375px  0 px changed (0.00%)\npass  undo    1280px  0 px changed (0.00%)\n--- impl-b\nFAIL  default  375px  1667 px changed (2.22%)\nFAIL  default 1280px  2103 px changed (0.82%)\nFAIL  long     375px  3363 px changed (4.48%)\nFAIL  long    1280px  4686 px changed (1.83%)\nFAIL  undo     375px  1799 px changed (2.40%)\nFAIL  undo    1280px  2235 px changed (0.87%)\n```\n\nRead it as a pixel count and B looks worse: 6 failures against A's 2. Read it as a reviewer and it's the other way round:\n\nSo the reviewer writes:\n\n```\nT-512 verdict, impl-a\n  visual: default pass   visual: long FAIL   visual: undo pass\n  screen reader: not checked   keyboard: not checked\n  Decision: send back. Long message must clamp at 2 lines (rule in\n  reference). Diff output attached.\n\nT-512 verdict, impl-b\n  visual: default/long/undo FAIL in check.mjs, 0.8% to 4.5%\n  Cause: 14 px vertical padding vs 12 px in reference\n  screen reader: ______   keyboard: ______   (reviewer runs these)\n  Decision: if both manual rows pass, accept with noted deviation:\n  14 px is the shared spacing token, so design updates the reference.\n```\n\nThe manual rows for B are left blank on purpose. I didn't run a screen reader or a keyboard pass for this post, and a verdict template shouldn't arrive pre-filled. Two things here are the reason the verdict exists. \"Not checked\" is a valid entry: A was sent back before anyone checked the screen reader, and the verdict says so instead of leaving it blank. And \"accept with noted deviation\" records a human decision to overrule a failing check, with a reason someone can disagree with later.\n\nWagglet's [request-to-delivery workflow guide](https://wagglet.com/blog/wagglet-workflow-request-draft-ticket-delivery) treats a delivery and its acceptance as separate events with separate owners, which is the same split this template writes down for one UI task.", "url": "https://wpnews.pro/news/a-screenshot-diff-can-t-accept-a-ui-change-for-you-a-task-template-with-visual", "canonical_source": "https://dev.to/sam_novak_574b07811e18495/a-screenshot-diff-cant-accept-a-ui-change-for-you-a-task-template-with-visual-reference-test-37k6", "published_at": "2026-10-08 09:12:56+00:00", "updated_at": "2026-10-08 09:18:11.348386+00:00", "lang": "en", "topics": ["ai-agents", "developer-tools", "ai-tools"], "entities": ["pixelmatch", "headless Chrome"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/a-screenshot-diff-can-t-accept-a-ui-change-for-you-a-task-template-with-visual", "markdown": "https://wpnews.pro/news/a-screenshot-diff-can-t-accept-a-ui-change-for-you-a-task-template-with-visual.md", "text": "https://wpnews.pro/news/a-screenshot-diff-can-t-accept-a-ui-change-for-you-a-task-template-with-visual.txt", "jsonld": "https://wpnews.pro/news/a-screenshot-diff-can-t-accept-a-ui-change-for-you-a-task-template-with-visual.jsonld"}}