# How I Built an iPhone App in Four Days with Opus 5.5

> Source: <https://projectautonomy.substack.com/p/i-built-an-iphone-app-in-four-days>
> Published: 2026-09-29 18:57:26+00:00

I recently built a working iPhone app in four days with Claude Opus 5.5, the newest frontier model. The app is for my girlfriend’s reselling business: it identifies products from photos and estimates what they’ll resell for, based on recent listings on eBay, Mercari, and other marketplaces.

Frontier models keep getting smarter and cheaper every few months. Opus 5.5 [outperforms Fable 5.1](https://www.anthropic.com/claude-opus-5-5) and [costs less](https://venturebeat.com/technology/anthropic-releases-claude-opus-5-5-beating-fable-5-1-on-key-agentic-benchmarks-at-60-cheaper-api-price), and open-weight models like [GLM-5.3](https://huggingface.co/zai-org/GLM-5.3) and [DeepSeek v4.1 Flash](http://www.deepseek.com/en/news/deepseek-v4-1-flash/) keep the pressure on. But benchmarks only tell you so much. Here’s what it was like building an app with Opus 5.5, including the impressive $4 test it devised that took the app’s accuracy from about 80% to 99%.

# First, we plan.

I started with a two-hour planning session with Claude. I opened with a long rant about my vision for the app, and at regular checkpoints I had Claude write what we’d agreed on into a plan document and keep it organized. We went through every major decision, starting with the basic user flows and working toward the technical side, like which services to rely on and how to keep running costs low.

That plan document is what I gave Claude Code to build from. A plan keeps an agent working from your rules and requirements instead of inventing its own as it goes. It also helped when Claude handed isolated tasks off to subagents: each one could compare its work to the grand plan.

Here’s roughly what went into the plan:

- The big idea: what the app does and who it’s for.
- User flows: how someone uses the app from the moment they open it.
- Design considerations: how I wanted the app to look, feel, and work.
- Technical decisions: I’m a software engineer, so I have opinions about how things get built, but I tried not to spend much time here. Mostly, I picked which third-party services to rely on. It’s best for you to choose how you spend your money, not the agent.

In the end, I settled on an Expo + React Native app with Supabase for the backend and database. I also pulled in a few other APIs, like Google Gemini and Jev, for the AI features. But this isn’t just an AI wrapper. It’s still 99% custom app logic.

Supabase works well with agents because it abstracts a lot of complex behavior behind a system that’s easy to understand. Since Supabase sets the rules, the agents made far fewer mistakes with configuration, migrations, and architecture. It’s also cheap enough that I don’t mind letting agents use it for testing. The free plan is so generous that I can keep using it until the app is done.

# Next, we build.

Once the plan was done, I handed it to Claude Code in Auto mode and said, “build this app following the plan. when I come back, I should be able to use the app without any issues.” Sometimes I just say shit because I think it will help. I really don’t know if it did here.

About 2 hours and 8,000 lines of code later, I had my girlfriend’s dream app built and ready to use. It looked nearly identical to what I’d described in my voice notes. I hadn’t said anything about theme or style, only structure and organization, and it went with a standard Apple look that I loved. Some of the more complex screens looked a little off, but that could wait until it was time to iterate.

Next, I needed to test it. I installed it on my iPhone and started snapping pictures of items to see if it recognized them. It did, mostly. Recognition worked, but the matching wasn’t tuned yet. It would pull up similar items instead of identical ones, and not consistently, which meant the price estimates were off. Still, I’d just skipped three to six months of part-time coding, so I wasn’t disappointed at all.

I was so excited with the early results that I sent Claude this:

It’s working, and it’s amazing, Claude. Genuinely, of all the things I’ve ever had you work on, this is the most impressive. It looks like an iOS app. It functions perfectly. It’s really well built, and it works on the first try without changes. You outdid yourself, Claude. Give yourself a pat on the back.

It did not function perfectly. I was just excited. I don’t normally anthropomorphize models, but this one earned it.

# Lastly, we iterate and improve.

If you’re happy being a meat vessel for AI, you can stop here. But the first version an agent hands you is a rough draft, and refining it is where you get give it your human element.

I started a voice recording and walked through the whole app, describing what I saw and what I wanted to see instead. I chose my words carefully, since the recording would be transcribed to text and had to make sense to an agent. I covered about 10 issues: layout changes, how interactions should behave, animations and sounds for key actions, and the general complaints a picky reviewer would have.

I sent the 5-minute recording straight to the agent, which transcribed it with tools on my computer, asked a few questions, and got to work. About an hour and a half later, I came back to a much better app. It finally looked the way I wanted, and the new features worked the way I’d described them.

The algorithm still had problems, though, and I didn’t understand exactly what was wrong until I took the app out and used it for real. I made another recording explaining what was going wrong, what I noticed, and exactly which data I was looking at when it happened.

The agent used my examples to find the exact bad results in the database and trace where the algorithm was making mistakes. Since the fix meant big changes to the core of the app, I had it plan first. I read through everything it proposed, then told it to go ahead. It worked for nearly 2 hours, making changes and testing the app itself to see whether the results improved.

Then I had the agent build its own benchmark: 33 real photos I’d taken, run through the search to measure how good the results were. That made the biggest difference of anything. Accuracy went from about 80% to 99%, and missed results dropped from about 19% to 1%. Thirty-three photos is a small test set, so I’m treating those numbers as a good sign rather than proof. The whole thing cost about $4 in API calls.

I’m still taking it into the real world, finding weak spots, and feeding them back to the agent. That loop of using it, noticing what to change and describing it clearly is where most of the work happens now. It’s the part I’d tell anyone building with these models not to skip.

Through this process of iterating, I got features added after the first build like:

- Grid and list views of scanned products.
- Quick filter buttons to see what products came back with exact matches, similar matches, or no matches.
- A like and dislike button on the result pages, so the search algorithm can be improved with time and testing.
- Streamlined product details (AI likes to repeat information and yap a lot; cutting and reimagining can help a lot with this).

# Conclusion.

This app started as a rant in a planning session. Four days later it was on my phone, and it gets a little better every time I take it out and use it.

The first build got me excited, but the version I have now works exactly the way I imagined. The model can write 20,000 lines of code while I’m away from my desk. It doesn’t have the capability to know what the app needs to be. That’s why it’s my job to test the app, and think about how others will feel while using it.

Intentional design is the new era.

One caveat: I wouldn’t build an app this way without some coding experience. Working straight from [Claude Code (limited free week codes)](https://claude.ai/referral/6MVV3GDJ5Q) and a Git repo leaves a lot of room for problems. If you’re not technical, tools like [Lovable](https://lovable.dev/invite/F2DKZDC), [Replit](https://replit.com/refer/fivemoreminix), and [Base44](https://app.base44.com/register?ref=XUYPWAYDPFYL1SLU&source=referral_program) (referral codes) are a better place to start. They give the agent its own instructions, pick a standard set of tools for you, and give you a real interface with one-click run buttons. With Claude Code you set all of that up yourself, but it costs less and you control everything.

If you want to see what I make next, subscribe. I’ll keep writing about what these models can and can’t do as I explore them.
