{"slug": "why-screenshot-diffing-ai-built-uis-thrashes-and-how-geometry-fixes-it", "title": "Why screenshot-diffing AI-built UIs thrashes, and how geometry fixes it", "summary": "A developer built designfit, an open-source MCP server and Claude Code skill that validates AI-generated front-end code against Figma frames by measuring design tokens and element geometry instead of diffing screenshots. The tool compares colors with CIEDE2000 perceptual distance and positions relative to the screen root, returning a machine-actionable fix-list; a demo run on a 360-node Figma frame scored three build iterations from 87 to 100.", "body_md": "If you've ever pointed a coding agent at a Figma frame and said \"make it match,\" you know the failure mode. The agent builds something close. You, or a tool, compare it to the design. It gets told \"still wrong,\" tweaks, compares again, and somehow it's *still* wrong. Forever. The agent isn't broken. The comparison is.\n\nThis post is about why the obvious comparison, diffing screenshots, is the wrong primitive for closing the loop with an AI agent, and what to use instead. The short version: don't diff pixels, measure geometry and tokens. That's the idea behind [designfit](https://github.com/as9978/designfit), an open-source MCP server and Claude Code skill I built. It extracts the spec from a Figma frame, validates the rendered front-end against it, and hands the agent a machine-actionable fix-list instead of an image.\n\n**▶ [Watch the 35-second demo](https://github.com/as9978/designfit#readme)**: a real run on a 360-node Figma frame, scoring three build iterations from 87 to 100.\n\nScreenshot-diffing means rendering your implementation, exporting the Figma frame, and computing a per-pixel difference. It's a great regression tool for a UI that's already correct. It's a terrible *convergence* tool for a UI that an agent is actively building.\n\nA rendered browser screenshot is full of differences that aren't mistakes. Anti-aliasing paints the edges of text and rounded corners with intermediate colors that depend on the exact sub-pixel position of each glyph. Font hinting and the platform's rasterizer shift things by fractions of a pixel. Sub-pixel layout rounding nudges a box by less than one device pixel. None of that is a design error, but every one of those pixels shows up in the diff as \"different.\"\n\nSo the agent gets a signal that says \"still wrong,\" with no way to tell a genuinely wrong color from a row of anti-aliased pixels along a letter's edge. It keeps editing. Because the diff is sensitive to noise the agent can't control, the edits don't reliably drive the number to zero. The score wobbles, the agent thrashes, and you burn tokens and time without getting closer to \"matches the design.\"\n\nThe deeper issue is determinism. A useful feedback loop needs **same input, same output**. A pixel diff doesn't have that property across machines, font stacks, or even re-renders, because its inputs include the rasterizer's noise. If the measurement isn't deterministic, the loop can't converge.\n\nA designer reviewing an implementation against a mockup doesn't overlay two images and hunt for differing pixels. They catch a small, structured set of things:\n\nThat's it. They're comparing a handful of meaningful, named properties against the design intent. Every one of those checks can be measured deterministically: a color is a color, a width is a number of pixels, a position is a coordinate, and an element is present or it isn't.\n\nSo the right primitive isn't \"how many pixels differ.\" It's \"which of the design's declared properties does the implementation violate.\" That set is small, stable, and machine-actionable, which is exactly what an agent needs to fix things instead of flailing.\n\ndesignfit validates two kinds of things, and nothing else.\n\n**Design tokens.** The resolved style values a design specifies: `fill`, `color`, `fontFamily`, `fontSize`, `fontWeight`, `lineHeight`, `letterSpacing`, `borderRadius`, `borderColor`, `borderWidth`, `opacity`. Colors are compared with **CIEDE2000 (ΔE)**, a perceptual color distance, so \"imperceptibly different\" doesn't read as a failure. Numeric tokens are compared in pixels.\n\n**Geometry.** Each element's box (`x`, `y`, `width`, `height`), compared **relative to the screen root** rather than to absolute viewport coordinates. A correctly built screen that happens to be centered or offset still passes, because every position is normalized to the root's origin first. You're measuring layout, not where the whole page landed.\n\nPlus **presence**: is each element the design expects in the DOM, and is anything tagged that the design doesn't know about.\n\nEvery comparison has an **explicit tolerance**, so sub-pixel and imperceptible-color noise never registers. The defaults:\n\n| Property | Tolerance | \n|---|---|\n| Geometry position / size | ±2 px | \n| Color | ΔE ≤ 2 | \n| Font size | ±1 px | \n| Line height | ±2 px | \n| Letter spacing | ±0.5 px | \n| Border width / radius | ±1 px | \n\nYou can loosen or tighten any of them per project.\n\nBecause the inputs are computed style values and bounding-box measurements rather than a rasterized image, the comparison is deterministic: same implementation, same design, same result, every run. That's what makes the loop converge. It also lets the agent stop early: if the score doesn't improve across two runs, the last edit didn't change anything designfit measures, so there's nothing left to thrash on.\n\ndesignfit exposes two MCP tools. `designfit_extract` turns a Figma frame into the input for validation. `designfit_validate` takes that input plus the running URL and returns `{ pass, score, violations, unmapped }`. A Claude Code skill drives the agent through the cycle:\n\n`designfit_extract` the frame's Figma link. It returns the design tree (each node with its frame and tokens), a component map, and the viewport. Figma variables and published styles bound to a property are recorded as that property's token source.`data-designfit-id=\"<figmaNodeId>\"`.` pass` is `true`, or stop and report if the score stalls.\nIn the first version, step 1 didn't exist. The agent read the frame through Figma's own MCP server and assembled the spec by hand. That was the most error-prone step in the loop: a mistyped frame or a forgotten token source produces a spec that disagrees with the design, and then a perfectly deterministic comparison deterministically checks the wrong thing.\n\n`designfit_extract` replaces that with code. It reads the frame from Figma's REST API with a personal access token (the Claude Code plugin asks for it once), or from `/nodes` JSON you paste. Like the rest of the tool, it's deterministic: same node JSON in, same spec out. Hidden nodes are skipped, a frame made only of vectors (an icon) becomes one leaf, and gradients, images and effects are left out.\n\nThe demo above is a real run, and it was more instructive than I expected.\n\nThe frame is a dense dark dashboard: 360 nodes once extracted, which took about 1.5 seconds. I validated 25 tagged elements across three iterations:\n\nTwo of those geometry errors weren't sizing mistakes at all. A section was 49 px tall in the build and 542 px in Figma; a table sat 73 px higher than designed. Both elements were **tagged on the wrong node**: the tag was on the section's header row instead of the section, and on the table's wrapper instead of the table. The original, hand-assembled spec had been built around those wrong tags, so it never noticed. Reading the frame straight from Figma did.\n\nRunning extract on a real frame also caught two bugs in designfit itself, both fixed in 0.2.1. Figma's per-side stroke weights were ignored, so a bottom-only border was expected on top. And CSS `letter-spacing: normal` was read as \"no value\", so every Figma text node with `letterSpacing: 0` raised a warning.\n\nEach violation names the component, the check, the property, what was expected (with the token's source when there is one), the actual value, a delta, and a fix hint:\n\n```\n{\n  \"component\": \"Button/Primary#btn\",\n  \"check\": \"token\",\n  \"property\": \"fill\",\n  \"expected\": { \"value\": \"#1d4ed8\", \"source\": \"color/primary\" },\n  \"actual\": { \"value\": \"#2b6cf0\" },\n  \"delta\": \"ΔE 6.4 (#2b6cf0 vs #1d4ed8)\",\n  \"severity\": \"error\",\n  \"fixHint\": \"use color/primary (#1d4ed8)\"\n}\n```\n\nThe `source` field is also the enforcement switch. A property bound to a Figma variable or published style is enforced: a mismatch is an `error` that fails the run and lowers the score. A hardcoded value with no token behind it is only a `warn`: it shows up in the fix-list so you don't miss it, but it never fails the run, because the only possible fix is another magic number. Geometry is always enforced. `pass` means zero errors.\n\ndesignfit 0.2 extracts the spec from Figma, then runs token, geometry and presence checks against **one viewport**: the breakpoint the frame was designed at. That's the whole product today, on purpose.\n\nExplicitly *roadmap, not shipped*: spacing as named tokens (wrong padding and gaps are caught today, but only as geometry deltas), validation across multiple breakpoints, and an advisory visual-model layer for the judgments geometry and tokens can't capture. Nothing fuzzy will ever sit in the pass/fail path. The bet is that most \"make it match the design\" thrash comes from wrong colors, wrong sizes and missing elements, and that deterministic measurement alone can kill it.\n\nIt's free and MIT-licensed. In Claude Code:\n\n```\n/plugin marketplace add as9978/designfit\n/plugin install designfit@designfit\n```\n\nIn any other MCP client, `npm install -g designfit`. Either way, run `npx playwright install chromium` once. Then paste a Figma link and let the agent close the loop on something it can measure.\n\nI'm the maker and building this solo. I'd like to hear where the tokens-and-geometry model breaks down on your frames: [open an issue](https://github.com/as9978/designfit/issues) with a case it missed, or one it flagged that wasn't real.\n\n**Validate AI-built front-ends against their Figma design — without the screenshot-diff thrash.**\n\nA real run on a 360-node Figma frame: `designfit_extract` reads the frame from its link, then `designfit_validate` scores three build iterations, 87 to 89 to 100 `pass`. No screenshot diffing anywhere in it.\n\n`designfit` is an MCP server + Claude Code skill that checks a rendered implementation against its Figma design and hands the coding agent a machine-actionable fix-list. It compares **design tokens** and **geometry** (element boxes relative to the screen root) — not raw pixels — so font-rendering noise never makes the agent oscillate. Deterministic in, deterministic out.\n\nScreenshot-diffing an AI-built UI against a Figma frame thrashes: anti-aliasing and sub-pixel shifts read as \"still wrong,\" so the agent fixes forever. designfit compares what a designer actually catches — wrong colors, wrong sizes, misalignment, missing elements — as **deterministic measurements with**…", "url": "https://wpnews.pro/news/why-screenshot-diffing-ai-built-uis-thrashes-and-how-geometry-fixes-it", "canonical_source": "https://dev.to/as9978/why-screenshot-diffing-ai-built-uis-thrashes-and-how-geometry-fixes-it-5967", "published_at": "2026-10-08 19:36:08+00:00", "updated_at": "2026-10-08 19:49:02.845424+00:00", "lang": "en", "topics": ["ai-agents", "agent-protocols", "developer-tools", "ai-tools"], "entities": ["designfit", "Figma", "Claude Code", "MCP"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/why-screenshot-diffing-ai-built-uis-thrashes-and-how-geometry-fixes-it", "markdown": "https://wpnews.pro/news/why-screenshot-diffing-ai-built-uis-thrashes-and-how-geometry-fixes-it.md", "text": "https://wpnews.pro/news/why-screenshot-diffing-ai-built-uis-thrashes-and-how-geometry-fixes-it.txt", "jsonld": "https://wpnews.pro/news/why-screenshot-diffing-ai-built-uis-thrashes-and-how-geometry-fixes-it.jsonld"}}