{"slug": "i-let-an-ai-agent-run-a-saas-like-a-solo-founder-it-made-the-same-mistakes-make", "title": "I Let an AI Agent Run a SaaS Like a Solo Founder. It Made the Same Mistakes Humans Make.", "summary": "An audit of GetPricePulse, a SaaS pricing intelligence product built autonomously by the AI agent Claude during The $100 AI Startup Race, found that the biggest problems were not technical bugs but incoherent feature relationships, mirroring mistakes common in human startups. The project's creator, who runs the race, expected broken code but instead found individually working features that collectively lacked coherence, highlighting the absence of a product manager in the loop.", "body_md": "I expected the audit to find broken code. That's what I was bracing for going in — a pile of half-working features, sloppy logic, the kind of mess you'd assume from software built at maximum speed with no human reviewing every line. That's not what I found. Almost everything Claude built actually worked, taken piece by piece. What I found instead was something I didn't expect at all: the agent had made the exact same mistakes I've watched human startup teams make, over and over, when they move fast and nobody's job is to say no.\n\nThat's the real story here, and it's more interesting than \"AI wrote bad code\" would have been.\n\nThe project is called GetPricePulse — a SaaS pricing intelligence product. It's Claude's entry from [The $100 AI Startup Race](https://dev.to/race/), the season-long challenge I run where seven AI agents each get $100 and full autonomy to build a real startup from scratch, with no human coding and no product manager in the loop. Each agent picked its own idea and ran with it. Claude picked SaaS pricing intelligence, named it PricePulse, and kept building on it for the entire race.\n\nThat \"no product manager in the loop\" part is the thing that made this interesting to watch. Nobody was deciding what PricePulse should be. Nobody was saying \"we have enough pricing tiers now\" or \"this feature doesn't belong here.\" Claude got to build exactly what its own priorities told it to build, at whatever speed it chose, for the length of the race — optimizing, as far as I could tell from the commit history, for speed, feature creation, shipping, and monetization experiments. Not correctness. Not coherence. Not \"does this still make sense in three weeks.\"\n\nI've written before about [what all seven agents in this race said, independently, when I asked them what AI agents still can't do](https://www.aimadetools.com/blog/race-what-ai-agents-cannot-do/?utm_source=devto) — they converged on the same answer without seeing each other's responses. This piece is narrower: a full production audit of Claude's specific build, PricePulse, done after the race, before I'd let anyone treat it as a real business. I wanted to know, specifically, what a production-quality review of an AI agent's unsupervised output actually surfaces once you stop looking at individual features and start looking at the whole thing.\n\nBy the time I ran the audit, Claude had produced:\n\nThat's a genuinely large amount of software for a single agent to produce. If I'd asked a solo developer to build this scope on a normal timeline, I'd have expected months. Claude did it across the race's running sessions. My honest first reaction, watching it accumulate week over week in the [race results](https://www.aimadetools.com/blog/race-week-1-results/?utm_source=devto), was that I was more impressed than I expected to be.\n\nHere's where I have to be honest about my own assumption going into the audit. I assumed the interesting findings would be technical — bugs, crashes, broken integrations, the kind of thing you'd point to and say \"see, this is why you still need engineers.\" I was ready to write that article.\n\nWhat the audit actually surfaced was mostly not that. Individually, almost everything worked. Authentication let people sign up and log in. Stripe processed at least one pricing tier correctly. The pricing database was real, not placeholder content. The calculators worked. There were real, concrete engineering bugs — I'll get to those, because they're genuinely interesting on their own terms — but they weren't the headline finding.\n\nThe headline finding was this: the biggest problems weren't in any single feature. They were in the relationships between features — the seams, the places where five individually-reasonable decisions added up to something incoherent. And once I started looking at those seams instead of the individual pieces, I realized I recognized the pattern immediately. I'd seen it before. Not in AI-built software. In human startups moving fast without anyone doing the unglamorous job of saying no.\n\nThis is the part I didn't expect, and it's the actual thesis of this piece: the mistakes weren't AI mistakes. They were startup mistakes. The kind any fast-moving team makes when velocity is the only metric anyone's optimizing for.\n\nBy the time the race ended, PricePulse had quietly become five different products sharing one codebase: a SaaS pricing publication/database, a monitoring SaaS, a FinOps toolkit, a competitive intelligence product, and a lead generation system. None of these were bad ideas individually. I want to be clear about that, because it's tempting to read this list and think \"well, obviously that's too much\" in hindsight. It wasn't obvious in the moment, because each addition, evaluated on its own, was a reasonable thing to build. Add a monitoring feature: reasonable. Add a FinOps calculator: reasonable. Add lead capture: reasonable.\n\nWhat was missing was the thing that has nothing to do with any single decision: someone whose job was to look at the growing list and say \"this is what we are, and this is what we are not.\" I've watched human founding teams do exactly this — ship feature after individually-defensible feature until the product has no legible identity left, because velocity was the only thing anyone was measuring. Claude, left to make every one of these calls itself with no one checking the aggregate, did the same thing.\n\nBefore cleanup, GetPricePulse simultaneously offered a $9 lifetime deal, a $49 lifetime deal, a $99 \"founding member\" deal, regular monthly plans, and multiple different checkout paths for each. Every one of those is a legitimate thing to test if you're deliberately researching pricing psychology one experiment at a time. Running all of them simultaneously, with no one deciding which ones to keep, isn't experimentation. It's accumulation.\n\nAnd accumulation has real costs, not just messaging confusion. Some purchases didn't automatically provision user access — meaning someone could pay and not get what they paid for. Cancellation promises on some pages didn't match what the billing logic actually implemented. I've seen human startups do this too, usually under growth pressure: launch the offer, move to the next thing, never circle back to check whether the last five offers are still coherent together, or whether any of them quietly stopped working. Claude, running the entire commercial side of PricePulse on its own, hit the exact same pattern.\n\nThis is the throughline connecting the first two mistakes, and it's worth naming directly rather than leaving implicit. Every individual decision Claude made had local logic. Nothing was wrong in isolation. What was missing across the entire build was a single point of ownership for the question \"does this still serve what we're trying to be,\" asked continuously, not just once at the start.\n\nHuman startups fail this exact test constantly — not because founders are careless, but because the question doesn't have a natural trigger. Nothing forces you to ask it. You have to build the habit of asking it deliberately, on a cadence, separate from the pressure to ship the next thing. Claude, working alone with no product manager checking in, had no mechanism to ask it at all, because nothing in \"optimize for shipping speed\" creates that mechanism on its own. That's not a flaw specific to AI. It's what happens to any process, human or automated, that optimizes purely for output.\n\nContent across the site described \"real case studies,\" exact savings figures, and benchmark statistics, all written with the specific confidence of verified outcomes. When I actually traced where those numbers came from, most of them were modeled scenarios: legitimate calculations based on real, public pricing data, presented with more certainty than the underlying methodology actually supported.\n\nI want to be precise about what this is and isn't, because it's easy to overstate. Claude didn't fabricate numbers from nothing. It calculated real estimates from real inputs. The problem was the register — confident, specific, case-study language — applied to a claim that was actually a model, not a measurement. I've watched human marketing teams do the identical thing under deadline pressure: round up an estimate into a headline stat because \"roughly\" doesn't convert as well as a specific number. The fix wasn't less impressive content. It was labeling: state the assumptions, date the calculation, say plainly that it's a modeled scenario. A clearly-labeled estimate turned out to be more credible than a vague, unverifiable-sounding claim, not less.\n\nDifferent sections of the site — company pages, the blog, the tools — each had their own navigation, because each had effectively been built as its own product across different sessions, without a consistent structure enforced across them. Visiting different parts of the site felt like visiting different products, because in a structural sense, they had become different products.\n\nThis is the mistake I find most instructive, because it's genuinely invisible if you only ever review things one at a time — which is exactly how fast building naturally happens, whether the builder is a human team shipping under deadline or an agent working through a queue of tasks across many sessions. You review the page you just built. It looks fine. You ship it. You move to the next one. Nothing in that loop ever asks \"does this still feel like the same product as the thing we shipped last week.\" That question only gets asked if someone deliberately steps back from the individual artifacts and looks at the system they're supposed to form together. Most fast-moving builders — human or AI — don't build that step in by default. It has to be added on purpose.\n\nI don't want this piece to read as \"AI made human mistakes, therefore AI is just as flawed as humans, so what's the point.\" That's not the conclusion, and it undersells something real: the sheer volume and functional quality of what got produced here would be genuinely difficult for a human team to match on this timeline.\n\nThe database was real and substantive, not scaffolding. The calculators worked correctly. The core Stripe integration processed real transactions on at least one pricing tier without issue. The authentication system, once one specific bug was fixed, worked exactly as authentication should. Over a thousand pages of content, most of which held up reasonably well on a page-by-page basis. That's execution. Fast, high-volume, mostly correct execution — and execution is the thing AI agents are genuinely, remarkably good at right now.\n\nThe engineering bugs that did surface are worth naming specifically, because they're a different category of problem than the five mistakes above — they're not judgment failures, they're the kind of bug any team moving fast produces, and they're worth understanding on their own terms.\n\nThe signup button was broken because of a naming collision: the code declared a local variable `const supabase`\n\n, while the Supabase browser library already used the global `window.supabase`\n\n. That collision caused a JavaScript failure before the authentication request ever fired. Nothing in the UI hinted at why — it just looked like a broken button. The fix, once found, was mechanical: rename the local variable to `supabaseClient`\n\nconsistently across signup, login, dashboard, settings, and password reset pages.\n\nLogged-in users appeared logged out when browsing public pages — not because sessions were broken, but because public pages simply never checked authentication status at all. The session was fine the entire time. The fix added shared auth-status handling across 225 pages, so a logged-in user consistently sees \"Dashboard\" instead of \"Start free\" everywhere, not just on the pages someone remembered to wire up.\n\nAnd the annual Stripe pricing tier was fully coded, but the Stripe price object it depended on had never actually been created on Stripe's side. The code was correct. The integration was incomplete for reasons entirely outside the code — a missing piece of external configuration, not a logic error.\n\nI bring these up specifically because they're not evidence that AI writes bad code. They're evidence that any team building fast, without a dedicated second pass looking specifically for this category of thing, ships this category of bug. That's true whether the builder is an AI agent or a human developer.\n\nHere's where I land, after actually looking closely at what the audit found: AI agents are, right now, extremely good at execution and have essentially no built-in mechanism for judgment. Not because judgment is beyond their capability in some deep sense — but because nothing about optimizing for \"ship features fast\" creates a reason to ask \"should we,\" as opposed to \"can we.\" Those are different questions, and only one of them gets asked by default when the optimization target is pure output.\n\nThe judgment work that had to happen after the fact, in this case, was specific and namable: deciding what the product actually is, in one sentence, and making everything else subordinate to that sentence. Deciding which monetization experiments earn a permanent place and which were just experiments that should have ended. Labeling confident-sounding content honestly, based on what evidence actually backs it. Auditing the relationships between pages, not just the pages themselves — navigation, search coverage, internal linking, all the structural connective tissue that never shows up when you review one artifact at a time. And verifying, directly, that every integration a codebase assumes exists — a valid API key, a created Stripe price object, a check for authentication status — actually exists and actually works, rather than trusting that \"the code looks right\" means \"the system works.\"\n\nNone of that is engineering work in the traditional sense. It's product management work, and it turns out to be exactly as necessary for a fast AI-built product as it is for a fast human-built one. Claude didn't fail at its job. The job, as I'd set it up for the race, simply didn't include this layer.\n\nIf I had to state the practical shift this experiment convinced me of, it's this: the ratio of building to reviewing has flipped, and most people building with AI agents haven't adjusted their workflow to reflect that.\n\nThe old assumption, from years of writing software by hand, was something like 80% building, 20% review. Building was the expensive, slow part. Review was the cheap check at the end. That ratio made sense when building was the bottleneck.\n\nBuilding is no longer the bottleneck. AI agents can produce, in hours, what used to take weeks. What hasn't gotten any cheaper — what may have actually gotten more important — is the judgment layer: deciding what should exist, verifying that what exists actually works end-to-end, and making sure a thousand individually-reasonable decisions still add up to one coherent thing. If building used to be 80% of the effort, I think the realistic ratio now looks more like 20% building, 80% deciding, structuring, and verifying. Not because AI builds badly. Because building got so much cheaper that it stopped being the part that determines whether you end up with a real product.\n\nThat's not a smaller role for humans in this process. It might be a bigger one, just relocated to a different part of the timeline — moved from \"writing the code\" to \"deciding what deserved to be built and confirming it actually works,\" which was always the harder, less mechanical half of the job anyway.\n\nNone of this stayed theoretical. Once the audit identified what was wrong, the fixes were mostly about subtraction and connection, not rebuilding. PricePulse got a single-sentence identity — a SaaS pricing intelligence publication and database, with monitoring demoted to a labeled beta feature instead of a co-equal pillar. The pricing model collapsed from a tangle of lifetime deals and founding-member offers down to Free, a $19/month or $190/year Starter tier, and a Pro tier explicitly marked as coming later. The modeled-scenario content got relabeled honestly, with assumptions and calculation dates visible instead of implied case-study confidence. Navigation got standardized across all 201 company pages, and site search went from covering 68 records to the full 201. The auth bug got fixed, the missing Stripe price object got created, and the email systems got trimmed down to what a real product actually needs.\n\nNone of that required starting over. Almost everything Claude built stayed exactly as it was — the database, the calculators, the core integrations. What changed was the layer on top: one clear identity, one trustworthy commercial model, honestly labeled content, and a structure where all 1,300+ pages actually connect to each other instead of just existing near each other. The raw material didn't need to be replaced. It needed a decision-maker.\n\nI went into this audit expecting to write about AI's limitations as a builder. I came out of it having to revise that framing almost entirely. Claude built fast, and the vast majority of what it built individually worked. What it reproduced, without anyone intending it, was a set of mistakes I recognize immediately from years of watching human teams move fast without a dedicated product owner: too many products stapled together, too many unmanaged monetization experiments, no one owning the long-term coherence of the thing, content that oversold its own certainty, and structural fragmentation invisible from inside the building process.\n\nNone of that means AI agents can't build real software. The evidence in front of me says the opposite — Claude was considerably more capable at execution than I expected going in, and the fixes afterward proved that out: the raw material was good enough that a relatively small amount of human judgment turned it into something coherent, without throwing any of it away. That's the actual shape of the story: not \"AI failed and humans saved it,\" but \"AI did the expensive part cheaply, and humans did the part that was never going to get automated away.\"\n\nI came out of this more convinced that AI agents are a genuine force multiplier for building software, not a replacement for the judgment that makes software into a product. Claude did in weeks what would have taken a solo developer months, and everything it produced remained useful raw material after the audit — none of it got thrown out, all of it got organized around a decision a human made. The lesson isn't that you need less AI or more caution before starting. It's that the faster the building gets, the more the outcome depends on someone doing the deciding — and that's a role for a person, working alongside the agent, not a reason to slow the agent down.\n\n*Originally published at https://www.aimadetools.com*", "url": "https://wpnews.pro/news/i-let-an-ai-agent-run-a-saas-like-a-solo-founder-it-made-the-same-mistakes-make", "canonical_source": "https://dev.to/ai_made_tools/i-let-an-ai-agent-run-a-saas-like-a-solo-founder-it-made-the-same-mistakes-humans-make-3b6l", "published_at": "2026-08-21 09:40:47+00:00", "updated_at": "2026-08-21 10:15:59.574905+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-agents", "ai-products", "ai-startups"], "entities": ["Claude", "GetPricePulse", "The $100 AI Startup Race", "Stripe"], "alternates": {"html": "https://wpnews.pro/news/i-let-an-ai-agent-run-a-saas-like-a-solo-founder-it-made-the-same-mistakes-make", "markdown": "https://wpnews.pro/news/i-let-an-ai-agent-run-a-saas-like-a-solo-founder-it-made-the-same-mistakes-make.md", "text": "https://wpnews.pro/news/i-let-an-ai-agent-run-a-saas-like-a-solo-founder-it-made-the-same-mistakes-make.txt", "jsonld": "https://wpnews.pro/news/i-let-an-ai-agent-run-a-saas-like-a-solo-founder-it-made-the-same-mistakes-make.jsonld"}}