cd /news/ai-agents/build-an-agent-loop-a-small-model-ca… · home › topics › ai-agents › article
[ARTICLE · art-143359] src=builder.io ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Build an agent loop a small model can finish

Agent-Native used a pixel-diff back-pressure check to let one of OpenAI's smaller models, Luna, cut a broken slide import's pixel mismatch from 88% to 2% over one weekend, according to the company's writeup. The team built the check by rendering the original and converted files and comparing them with pixelmatch or ImageMagick's compare, after finding four bugs in the check itself, including a font-loading error that reported 17% off when the real difference was 0.19%. The writeup argues the loop is easy and the verifier is the hard part, since an agent that can measure every attempt does not need exceptional judgment about its own work.

by read11 min views1 publishedOct 1, 2026
Build an agent loop a small model can finish
Image: Builder (auto-discovered)

Agent-Native We got a small, cheap coding agent to finish a job that looked too big for it. Then we built a second loop from scratch, on a job small enough to try yourself, to see which parts actually mattered.

An agent loop keeps a coding agent working on a task, one attempt after another, until something says it's done. The loop is the easy part. Most of the work goes into that "something": a check the agent can run by itself that says whether each attempt got closer.

If you've tried a Ralph loop or /goal, you've already seen the part where the agent keeps going. Geoffrey Huntley calls the other half "back pressure": tests, builds, type checks, and anything else that can reject a bad attempt. Neither of our jobs came with a test suite, so we had to build the back pressure ourselves. The first job was fixing Figma and slide imports in Agent-Native, our free, open-source framework and collection of apps. One slide came through our importer so broken that 88% of its pixels didn't match the original. Over one weekend, one of OpenAI's smaller models got that down to 2%. Here's a walkthrough of that loop.

The first loop: fixing our imports #

Our imports started out rough. People wanted to bring Figma files into Agent-Native Design and decks into Agent-Native Slides. Diagrams came apart into loose shapes. Icons disappeared. One slide came through as a black rectangle. So we looked at the results, described what was wrong, asked the agent for a fix, and looked again.

That loop could only move as fast as a person could look at slides, and the agent only knew what somebody remembered to tell it. We were the verifier, typing "still wrong, the heading moved again" into a chat over and over.

Turning "looks right" into a number

Comparing visuals feels like a judgment call. That's hard to write as a unit test or a browser test.

But every imported file already came with a right answer: the original. So we rendered the original and the converted version and compared them with a pixel diff.

A pixel diff compares two images and flags every pixel that differs. Tools like pixelmatch and ImageMagick's compare do this out of the box. Our check gave the agent two things: a number, and an image with the mismatches marked in red.

The number told the agent whether a change helped. The red overlay told it where to look next.

Checking the check

An agent working against a number will believe the number, so the number had to be right. At first, it wasn't. Fonts that never loaded made one export look 17% off when the real difference was 0.19%. Our notes from those runs found four bugs like that in the check itself: "A fidelity number is a claim about the converter, so the harness has to be at least as trustworthy as the thing it grades."

We also had to decide what didn't count. Font rendering differences weren't worth chasing, so we told the agent to ignore them.

Then the check needed real examples. A few tidy samples would've let the agent pass, so we had it collect a wide range of community files across .fig files, the Figma API, PowerPoint, PDF, and Google Slides. Every change was judged against the whole set. Our notes record two fixes that "made the frame better and the corpus worse," and both were reverted.

Letting a small model grind

With the check in place, the loop needed a model that'd keep trying. It didn't have to be our best one. An agent that can measure every attempt doesn't need exceptional judgment about its own work. It needs to make a change, run the check, see whether the number went down, and try again.

We used Luna, one of OpenAI's smaller models, on its max setting through a ChatGPT Pro plan. Our rough estimate is that it gets close to Opus-level results for a tenth to a twentieth of the price, and on that plan it can run for days without hitting a limit. We left it running through the weekend. Diffs that started at 20%, 40%, even 50% came down to single digits, often fractions of a percent.

Import and export problems had been our most common feedback, with dozens of reports a day. After the loop, they stopped coming in, and the edge cases people found later were easy to fix. An agent told not to stop until it gets there will keep going, and so will /goal in Codex or Claude Code.

The second loop: rebuilding a page from screenshots #

The import loop took a weekend of Luna's time, but most of the hard thinking happened before Luna started. To see that part clearly, we built a second loop from scratch on a job small enough for anyone to run: rebuild the top of a blog post as a static HTML and CSS page, working only from screenshots.

The agent got screenshots at three screen widths, the fonts and images, and the page text. The check renders its page in a headless browser and diffs it against each screenshot.

This time, we split the work on purpose. A strong agent builds the loop. A small agent runs against it.

Building the check

We worked through this prompt with Claude:

Claude built the check, a one-page brief for Luna, and a set of pages to test the check with: the original rendered twice, pages we knew were bad, and pages we'd accept.

The check failed that test twice. The first version counted mismatched pixels as a share of the whole screenshot. The page has a dark background, so most pixels match even when there's nothing on the page. A blank page scored 3 to 5%. We changed the score to count only the visible content, and a blank page went up to 57% or more.

The second version was too strict. At the narrowest width, moving the whole page down by one pixel scored 31%. Nobody would notice a one-pixel shift, so we told the check to forgive it. After that, the scores lined up with what you'd see:

| Page | 390px wide | 1280px wide | | The original, rendered twice | 0% | 0% | | Everything moved 1px | 0% | 0% | | Title moved 3px | 9% | 9% | | Hero image missing | 13% | 34% | | Wrong font | 48% | 41% | | Blank page | 57% | 76% |

The pass line was 2% of the visible content at every width, about what you get from a button that's 4 pixels out of place.

We also held examples back. Luna worked from screenshots at 390, 768, and 1280 pixels wide. We kept 1024 and 1440 to ourselves, so we could tell whether it had solved the page or just learned the screenshots. That turned out to matter.

Running the loop

Luna got this prompt:

The prompt puts the check off-limits. An agent working hard to lower a number may find it easier to change the measurement than the code. Across four rounds, Luna never touched it.

After each round, we looked at what Luna made before we looked at the score. A visual recap makes that quick. Round 1. We gave Luna, on its high setting, a 90-minute budget. Its first attempt scored 41.5%, and 34 minutes later it had stalled at 16.7%. But its page looked almost identical to the original. When a loop stalls, it's tempting to blame the model and switch to a bigger one. That wouldn't have helped here. The check was wrong: our one-pixel tolerance only forgave whole-pixel shifts, so text that landed half a pixel off still counted. We fixed the tolerance, and Luna's page dropped to about 2%.

At 1024 pixels wide, a width Luna never saw, its page switched to the full menu too early and lost its left margin.

Round 2. We added a line to the brief saying the page had to hold up at every width in between. Luna passed the three checked widths in nine minutes, and the holdout failed the same way. It had read the line, but it had no way to check it, so it fixed what the check could see. Anything that mattered had to be in the check, not just the brief.

Round 3. We moved 1024 and 1440 into the check. Once we'd used those widths to make decisions, they weren't really held back anymore, so we set aside three new widths. Luna passed all five checked widths in five minutes and failed all three new ones. That was the fastest pass of all four rounds and the worst result. A quick pass is a reason to look harder. This time, the reason was in the CSS. It showed the menu button only between 1024 and 1100 pixels, just enough to cover the one width it could see. It also nudged one paragraph's letter spacing by 0.03 pixels, so the text would wrap the right way at 390.

Round 4. Cheng Lou built Pretext by checking it against the browser itself: "showing Claude Code and Codex the browsers ground truth, and have them measure & iterate against those at every significant container width, running over weeks." Our check needed every significant width too. We gave it 28 widths and held back five more. Luna took 59 minutes and 55 check runs to pass all 28, and three of the five held-back widths passed.

That's a page that looks right almost everywhere. Its stylesheet is another matter: 21 media queries, many covering only 25 to 50 pixels. A person would write a few fluid rules. The check measured pixels, so pixels are what Luna matched. If we cared how the code was written, and for real work we would, that would need a check of its own.

| Round | What we changed | Luna's time | Check runs | Checked widths | Held-back widths | | 1 | First run | 34 min | 63 | Stuck at 16.7% | 0 of 2 pass | | 2 | Fixed the half-pixel tolerance | 9 min | 15 | 3 of 3 pass | 0 of 2 pass | | 3 | Checked 5 widths, held back 3 new ones | 5 min | 6 | 5 of 5 pass | 0 of 3 pass | | 4 | Checked 28 widths, held back 5 new ones | 59 min | 55 | 28 of 28 pass | 3 of 5 pass |

Across four rounds, Luna spent under two hours on the grinding. Every jump between rounds came from changing the check, not the model. That's a lot of work from a cheap model that just needed better boundaries.

What to take from this #

Both loops taught us the same things. If you're building one, here's where we'd put the effort:

  • Find the right answer you already have. Most hard problems have something like that: the output of an old implementation, the browser's own behavior, examples from a spec, recorded production traffic, or a benchmark you want to beat.
  • Test the check before you trust it. Render the same thing twice and make sure it scores zero. Then make sure results you know are bad score badly and results you'd accept score well. Write down what doesn't count, too, or the agent will spend hours on noise.
  • Use real examples, and hold some back. Judge every change against the whole set. Held-back examples show whether the agent solved the problem or just learned the examples.
  • Put what matters in the check. The agent does what the number rewards.
  • Keep the check out of reach. Tell the agent the check is fixed, and ask it to stop and explain if it thinks the check is wrong.
  • Give the building to a strong agent and the grinding to a small one. The two prompts above work for other jobs, too. Swap in your task.
  • Look at the output before the score. A stall can mean the check is wrong, not the model. A fast pass can mean the agent found a shortcut.
  • Keep the check around. Turn your final scores into ceilings in a test, so a later change that makes things worse fails it.

The check doesn't have to be visual, either. P90 response times, error rates, user behavior you can verify, browser workflows you can automate, and plain unit tests all give an agent a number or a pass or fail it can work against without asking you. If you've built a factory that picks up bug reports, checks like these are how its agents know a fix is done.

So pick one thing you've been checking by hand. What would it take to measure it?

If you'd like to see where our loop ended up, import a Figma file into Agent-Native Design or a deck into Agent-Native Slides. Both are free and open source.

── more in #ai-agents 4 stories · sorted by recency
── more on @agent-native 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/build-an-agent-loop-…] indexed:0 read:11min 2026-10-01 · —