If you've ever pointed a coding agent at a Figma frame and said "make it match," you know the failure mode. The agent builds something close. You, or a tool, compare it to the design. It gets told "still wrong," tweaks, compares again, and somehow it's still wrong. Forever. The agent isn't broken. The comparison is.
This post is about why the obvious comparison, diffing screenshots, is the wrong primitive for closing the loop with an AI agent, and what to use instead. The short version: don't diff pixels, measure geometry and tokens. That's the idea behind designfit, an open-source MCP server and Claude Code skill I built. It extracts the spec from a Figma frame, validates the rendered front-end against it, and hands the agent a machine-actionable fix-list instead of an image.
▶ Watch the 35-second demo: a real run on a 360-node Figma frame, scoring three build iterations from 87 to 100.
Screenshot-diffing means rendering your implementation, exporting the Figma frame, and computing a per-pixel difference. It's a great regression tool for a UI that's already correct. It's a terrible convergence tool for a UI that an agent is actively building.
A rendered browser screenshot is full of differences that aren't mistakes. Anti-aliasing paints the edges of text and rounded corners with intermediate colors that depend on the exact sub-pixel position of each glyph. Font hinting and the platform's rasterizer shift things by fractions of a pixel. Sub-pixel layout rounding nudges a box by less than one device pixel. None of that is a design error, but every one of those pixels shows up in the diff as "different."
So the agent gets a signal that says "still wrong," with no way to tell a genuinely wrong color from a row of anti-aliased pixels along a letter's edge. It keeps editing. Because the diff is sensitive to noise the agent can't control, the edits don't reliably drive the number to zero. The score wobbles, the agent thrashes, and you burn tokens and time without getting closer to "matches the design."
The deeper issue is determinism. A useful feedback loop needs same input, same output. A pixel diff doesn't have that property across machines, font stacks, or even re-renders, because its inputs include the rasterizer's noise. If the measurement isn't deterministic, the loop can't converge.
A designer reviewing an implementation against a mockup doesn't overlay two images and hunt for differing pixels. They catch a small, structured set of things:
That's it. They're comparing a handful of meaningful, named properties against the design intent. Every one of those checks can be measured deterministically: a color is a color, a width is a number of pixels, a position is a coordinate, and an element is present or it isn't.
So the right primitive isn't "how many pixels differ." It's "which of the design's declared properties does the implementation violate." That set is small, stable, and machine-actionable, which is exactly what an agent needs to fix things instead of flailing.
designfit validates two kinds of things, and nothing else.
Design tokens. The resolved style values a design specifies: fill, color, fontFamily, fontSize, fontWeight, lineHeight, letterSpacing, borderRadius, borderColor, borderWidth, opacity. Colors are compared with CIEDE2000 (ΔE), a perceptual color distance, so "imperceptibly different" doesn't read as a failure. Numeric tokens are compared in pixels.
Geometry. Each element's box (x, y, width, height), compared relative to the screen root rather than to absolute viewport coordinates. A correctly built screen that happens to be centered or offset still passes, because every position is normalized to the root's origin first. You're measuring layout, not where the whole page landed.
Plus presence: is each element the design expects in the DOM, and is anything tagged that the design doesn't know about.
Every comparison has an explicit tolerance, so sub-pixel and imperceptible-color noise never registers. The defaults:
| Property | Tolerance |
|---|---|
| Geometry position / size | ±2 px |
| Color | ΔE ≤ 2 |
| Font size | ±1 px |
| Line height | ±2 px |
| Letter spacing | ±0.5 px |
| Border width / radius | ±1 px |
You can loosen or tighten any of them per project.
Because the inputs are computed style values and bounding-box measurements rather than a rasterized image, the comparison is deterministic: same implementation, same design, same result, every run. That's what makes the loop converge. It also lets the agent stop early: if the score doesn't improve across two runs, the last edit didn't change anything designfit measures, so there's nothing left to thrash on.
designfit exposes two MCP tools. designfit_extract turns a Figma frame into the input for validation. designfit_validate takes that input plus the running URL and returns { pass, score, violations, unmapped }. A Claude Code skill drives the agent through the cycle:
designfit_extract the frame's Figma link. It returns the design tree (each node with its frame and tokens), a component map, and the viewport. Figma variables and published styles bound to a property are recorded as that property's token source.data-designfit-id="<figmaNodeId>". pass is true, or stop and report if the score stalls.
In the first version, step 1 didn't exist. The agent read the frame through Figma's own MCP server and assembled the spec by hand. That was the most error-prone step in the loop: a mistyped frame or a forgotten token source produces a spec that disagrees with the design, and then a perfectly deterministic comparison deterministically checks the wrong thing.
designfit_extract replaces that with code. It reads the frame from Figma's REST API with a personal access token (the Claude Code plugin asks for it once), or from /nodes JSON you paste. Like the rest of the tool, it's deterministic: same node JSON in, same spec out. Hidden nodes are skipped, a frame made only of vectors (an icon) becomes one leaf, and gradients, images and effects are left out.
The demo above is a real run, and it was more instructive than I expected.
The frame is a dense dark dashboard: 360 nodes once extracted, which took about 1.5 seconds. I validated 25 tagged elements across three iterations:
Two of those geometry errors weren't sizing mistakes at all. A section was 49 px tall in the build and 542 px in Figma; a table sat 73 px higher than designed. Both elements were tagged on the wrong node: the tag was on the section's header row instead of the section, and on the table's wrapper instead of the table. The original, hand-assembled spec had been built around those wrong tags, so it never noticed. Reading the frame straight from Figma did.
Running extract on a real frame also caught two bugs in designfit itself, both fixed in 0.2.1. Figma's per-side stroke weights were ignored, so a bottom-only border was expected on top. And CSS letter-spacing: normal was read as "no value", so every Figma text node with letterSpacing: 0 raised a warning.
Each violation names the component, the check, the property, what was expected (with the token's source when there is one), the actual value, a delta, and a fix hint:
{
"component": "Button/Primary#btn",
"check": "token",
"property": "fill",
"expected": { "value": "#1d4ed8", "source": "color/primary" },
"actual": { "value": "#2b6cf0" },
"delta": "ΔE 6.4 (#2b6cf0 vs #1d4ed8)",
"severity": "error",
"fixHint": "use color/primary (#1d4ed8)"
}
The source field is also the enforcement switch. A property bound to a Figma variable or published style is enforced: a mismatch is an error that fails the run and lowers the score. A hardcoded value with no token behind it is only a warn: it shows up in the fix-list so you don't miss it, but it never fails the run, because the only possible fix is another magic number. Geometry is always enforced. pass means zero errors.
designfit 0.2 extracts the spec from Figma, then runs token, geometry and presence checks against one viewport: the breakpoint the frame was designed at. That's the whole product today, on purpose.
Explicitly roadmap, not shipped: spacing as named tokens (wrong padding and gaps are caught today, but only as geometry deltas), validation across multiple breakpoints, and an advisory visual-model layer for the judgments geometry and tokens can't capture. Nothing fuzzy will ever sit in the pass/fail path. The bet is that most "make it match the design" thrash comes from wrong colors, wrong sizes and missing elements, and that deterministic measurement alone can kill it.
It's free and MIT-licensed. In Claude Code:
/plugin marketplace add as9978/designfit
/plugin install designfit@designfit
In any other MCP client, npm install -g designfit. Either way, run npx playwright install chromium once. Then paste a Figma link and let the agent close the loop on something it can measure.
I'm the maker and building this solo. I'd like to hear where the tokens-and-geometry model breaks down on your frames: open an issue with a case it missed, or one it flagged that wasn't real.
Validate AI-built front-ends against their Figma design — without the screenshot-diff thrash.
A real run on a 360-node Figma frame: designfit_extract reads the frame from its link, then designfit_validate scores three build iterations, 87 to 89 to 100 pass. No screenshot diffing anywhere in it.
designfit is an MCP server + Claude Code skill that checks a rendered implementation against its Figma design and hands the coding agent a machine-actionable fix-list. It compares design tokens and geometry (element boxes relative to the screen root) — not raw pixels — so font-rendering noise never makes the agent oscillate. Deterministic in, deterministic out.
Screenshot-diffing an AI-built UI against a Figma frame thrashes: anti-aliasing and sub-pixel shifts read as "still wrong," so the agent fixes forever. designfit compares what a designer actually catches — wrong colors, wrong sizes, misalignment, missing elements — as deterministic measurements with…