Why screenshot-diffing AI-built UIs thrashes, and how geometry fixes it A developer built designfit, an open-source MCP server and Claude Code skill that validates AI-generated front-end code against Figma frames by measuring design tokens and element geometry instead of diffing screenshots. The tool compares colors with CIEDE2000 perceptual distance and positions relative to the screen root, returning a machine-actionable fix-list; a demo run on a 360-node Figma frame scored three build iterations from 87 to 100. If you've ever pointed a coding agent at a Figma frame and said "make it match," you know the failure mode. The agent builds something close. You, or a tool, compare it to the design. It gets told "still wrong," tweaks, compares again, and somehow it's still wrong. Forever. The agent isn't broken. The comparison is. This post is about why the obvious comparison, diffing screenshots, is the wrong primitive for closing the loop with an AI agent, and what to use instead. The short version: don't diff pixels, measure geometry and tokens. That's the idea behind designfit https://github.com/as9978/designfit , an open-source MCP server and Claude Code skill I built. It extracts the spec from a Figma frame, validates the rendered front-end against it, and hands the agent a machine-actionable fix-list instead of an image. ▶ Watch the 35-second demo https://github.com/as9978/designfit readme : a real run on a 360-node Figma frame, scoring three build iterations from 87 to 100. Screenshot-diffing means rendering your implementation, exporting the Figma frame, and computing a per-pixel difference. It's a great regression tool for a UI that's already correct. It's a terrible convergence tool for a UI that an agent is actively building. A rendered browser screenshot is full of differences that aren't mistakes. Anti-aliasing paints the edges of text and rounded corners with intermediate colors that depend on the exact sub-pixel position of each glyph. Font hinting and the platform's rasterizer shift things by fractions of a pixel. Sub-pixel layout rounding nudges a box by less than one device pixel. None of that is a design error, but every one of those pixels shows up in the diff as "different." So the agent gets a signal that says "still wrong," with no way to tell a genuinely wrong color from a row of anti-aliased pixels along a letter's edge. It keeps editing. Because the diff is sensitive to noise the agent can't control, the edits don't reliably drive the number to zero. The score wobbles, the agent thrashes, and you burn tokens and time without getting closer to "matches the design." The deeper issue is determinism. A useful feedback loop needs same input, same output . A pixel diff doesn't have that property across machines, font stacks, or even re-renders, because its inputs include the rasterizer's noise. If the measurement isn't deterministic, the loop can't converge. A designer reviewing an implementation against a mockup doesn't overlay two images and hunt for differing pixels. They catch a small, structured set of things: That's it. They're comparing a handful of meaningful, named properties against the design intent. Every one of those checks can be measured deterministically: a color is a color, a width is a number of pixels, a position is a coordinate, and an element is present or it isn't. So the right primitive isn't "how many pixels differ." It's "which of the design's declared properties does the implementation violate." That set is small, stable, and machine-actionable, which is exactly what an agent needs to fix things instead of flailing. designfit validates two kinds of things, and nothing else. Design tokens. The resolved style values a design specifies: fill , color , fontFamily , fontSize , fontWeight , lineHeight , letterSpacing , borderRadius , borderColor , borderWidth , opacity . Colors are compared with CIEDE2000 ΔE , a perceptual color distance, so "imperceptibly different" doesn't read as a failure. Numeric tokens are compared in pixels. Geometry. Each element's box x , y , width , height , compared relative to the screen root rather than to absolute viewport coordinates. A correctly built screen that happens to be centered or offset still passes, because every position is normalized to the root's origin first. You're measuring layout, not where the whole page landed. Plus presence : is each element the design expects in the DOM, and is anything tagged that the design doesn't know about. Every comparison has an explicit tolerance , so sub-pixel and imperceptible-color noise never registers. The defaults: | Property | Tolerance | |---|---| | Geometry position / size | ±2 px | | Color | ΔE ≤ 2 | | Font size | ±1 px | | Line height | ±2 px | | Letter spacing | ±0.5 px | | Border width / radius | ±1 px | You can loosen or tighten any of them per project. Because the inputs are computed style values and bounding-box measurements rather than a rasterized image, the comparison is deterministic: same implementation, same design, same result, every run. That's what makes the loop converge. It also lets the agent stop early: if the score doesn't improve across two runs, the last edit didn't change anything designfit measures, so there's nothing left to thrash on. designfit exposes two MCP tools. designfit extract turns a Figma frame into the input for validation. designfit validate takes that input plus the running URL and returns { pass, score, violations, unmapped } . A Claude Code skill drives the agent through the cycle: designfit extract the frame's Figma link. It returns the design tree each node with its frame and tokens , a component map, and the viewport. Figma variables and published styles bound to a property are recorded as that property's token source. data-designfit-id="