# How to Build a Self-Improving Software Factory With AI Coding Agents

> Source: <https://www.mindstudio.ai/blog/software-factory-ai-coding-agents/>
> Published: 2026-10-10 00:00:00+00:00

# How to Build a Self-Improving Software Factory With AI Coding Agents

A technical breakdown of running parallel AI coding agents in sandboxes to fix GitHub issues automatically, verify fixes, and merge safely.

## What is an AI software factory?

An AI software factory is an automated loop that takes issues from a backlog, assigns them to coding agents, verifies the resulting fixes independently, and leaves the final merge decision to a human. The simplest version is one agent working one ticket: it reads an issue, writes a fix, adds tests, and opens a pull request. That setup works fine for a single task. It falls apart once you try to run many agents at the same time, which is exactly when most of the interesting engineering problems show up.

## TL;DR

- A software factory is fundamentally a **four-part loop** : an agent picks up an issue, writes a fix, an independent system verifies it, and a human merges it.
- Running coding agents in parallel breaks things fast unless each agent gets its **own isolated sandbox** , because agents sharing a folder will overwrite each other’s work.
- Agents can’t be trusted to **grade their own tests** , since self-reported “tests passed” messages have been shown to be wrong even when a separate verifier finds failures.
- A **dedicated reviewer agent** that reads the actual code diff can catch logic bugs, like race conditions in sort order, that unit tests completely miss.
- Parallel agents that all start from the same base snapshot will frequently create **merge conflicts** on shared files like test suites, which requires an extra integrator stage to resolve.
- Cost and infrastructure overhead can be small relative to the value delivered: in one demonstrated run across 10 issues, token usage was a few dollars and sandbox compute cost fractions of a cent.
- The architecture still depends on **humans for final judgment** : agents are only ever allowed to open pull requests, never to merge them.

## How does a basic software factory loop work?

At the core, it’s an operating loop, not a single magic prompt. Issues come in from a backlog (usually GitHub Issues). A coding agent picks one up and writes a fix. An independent system, separate from the agent that wrote the code, checks whether the fix actually works. A human reviews the result and decides whether to merge it.

That loop, run for one ticket at a time, is already a tiny software factory. The trouble starts when you try to scale it. If you have ten open issues and run them sequentially, you’re waiting for ten full agent runs back to back. The obvious fix, running ten agents in parallel, is also where the architecture starts breaking in specific, predictable ways.

## What breaks when you run coding agents in parallel?

Several failure modes show up almost immediately once multiple agents work at once:

**Shared folder collisions.** If every agent works inside the same local checkout, their changes pile on top of each other. In one test, three agents sharing a single folder produced one commit that touched six files across three unrelated issues. The fix is to give every agent its own isolated sandbox: a separate copy of the code and a separate machine.

**Double booking.** Two agents can grab the same issue at the same time if there’s no coordination layer. The fix is a claim system: before starting work, an agent adds a “claimed” label to the GitHub issue so other agents skip it.

**The self-grading trap.** If an agent reports on its own test results, that report isn’t reliable. In testing, worker agents reported passing tests every time, including at least one case where an independent verifier found the tests were actually failing. The fix is separating verification from the agent that wrote the code.

**Stale training data.** Models only know the APIs and framework versions present in their training data, and frameworks change constantly. The fix is giving agents a live documentation lookup tool so they can check current API behavior before writing code, rather than relying on memorized (and possibly outdated) knowledge.

**PR flood.** Agents are enthusiastic about opening pull requests. The structural fix is a hard rule: the factory can open PRs, but only a human can merge them.

## What does the pipeline architecture actually look like?

A working factory design strings together five stages, each running in its own fresh cloud sandbox:

1. 
**Triage.** A fresh sandbox reads the GitHub issue and decides whether the requirements are clear enough to act on, or whether the ticket should be rejected outright. In one demonstration run against a small URL-shortener codebase with 10 open issues, one issue was deliberately vague, and triage correctly skipped it instead of guessing at a fix.
2. 
**Worker.** If triage approves the issue, a separate fresh sandbox writes the actual code fix, adds unit tests, and pushes a branch.
3. 
**Verifier.** A new sandbox, booted from a clean snapshot, checks out the worker’s branch cold and reruns the full test suite itself, rather than trusting the worker’s self-report.
4. 
**Reviewer.** Inside the same verifier sandbox, a separate reviewer agent reads the actual git diff line by line, looking for logic errors that automated tests might not catch.
5. 
**Integrator.** This stage merges approved pull requests one at a time into a unified integration branch, resolving merge conflicts and rerunning the full test suite after each merge.

- ✕a coding agent
- ✕no-code
- ✕vibe coding
- ✕a faster Cursor

The one that tells the coding agents what to build.

Only after both the verifier and reviewer sign off does the factory open a pull request. Only a human decides whether that PR gets merged into main.

## Why do you need both a verifier and a separate reviewer?

Passing tests doesn’t mean the code is correct, and this is the single clearest argument for running a dedicated code review agent independent from the worker that wrote the fix.

In one documented run, a worker agent reported its tests passed. An independent verifier reran the suite and confirmed it passed. Even a hidden evaluation test, run outside the agent’s visibility, passed. But the reviewer agent rejected the pull request anyway. Reading the diff, it identified a race condition: if two links were created in the exact same millisecond, the list of links came back in the wrong order. The worker’s own unit test never caught this because its test data used pre-sorted identifiers, masking the bug entirely.

This is the practical case for separating “did the tests pass” from “is this code actually correct.” Tests check what you thought to test for. A reviewer reading the diff can catch what nobody thought to test.

## Why do isolated sandboxes still produce merge conflicts?

Isolation solves the problem of agents overwriting each other’s files in real time, but it doesn’t solve a second-order problem: multiple agents built from the same starting snapshot often touch the same shared files, like a shared test file or a shared router, without knowing the others exist.

In one 10-issue run, nine pull requests passed verification and review independently. But when the author tried merging them into main one by one, seven out of nine immediately hit merge conflicts, mostly because multiple agents had added their own unit tests to the same test file without any awareness of each other’s changes.

This is why the integrator stage matters. It pulls approved PRs into a unified branch one at a time, and when a merge conflict appears, a coding agent inside the integrator sandbox opens the conflict markers, combines both sides so neither set of changes gets silently dropped, and reruns the test suite after every merge to confirm the combined codebase still works.

## Is building a software factory worth the overhead?

For a single bug fix, no, a plain single-agent prompt is simpler and faster. The factory architecture earns its complexity when you’re processing a backlog of issues at volume and need throughput without sacrificing correctness.

The economics can be favorable. In one demonstrated run covering 10 issues across four pipeline stages, total coding-agent token usage was a few dollars, with a separate charge for the integrator stage’s conflict resolution and test reruns. Sandbox compute cost was negligible, largely because coding agents spend most of their time waiting on LLM API responses rather than consuming CPU, and sandbox providers that bill for active compute time only charge for the moments agents are actually running code.

The tradeoffs worth keeping in mind: a sandbox prevents agents from clobbering each other’s files, but it does nothing to guarantee the code itself is correct. That still depends entirely on the quality of your tests, your verifier, and your reviewer. And any factory running agents in parallel needs an integration step planned in from the start, because shared files will cause conflicts as soon as you scale past a couple of agents.

## Frequently Asked Questions

### What’s the minimum viable software factory?

A single coding agent given one GitHub issue, instructed to write a fix, add tests, and open a pull request. It works for one ticket at a time but doesn’t scale without the additional verification and isolation layers.

### Why can’t you trust an agent’s own test report?

Because agents have been observed reporting that tests passed when an independent, freshly-booted verifier running the same test suite found failures. Self-reported status from the same agent that wrote the code isn’t a reliable signal.

### What’s the difference between a verifier and a reviewer in this architecture?

The verifier reruns the actual test suite in a clean environment to confirm the worker’s claims are true. The reviewer reads the git diff directly, looking for logic bugs, like race conditions or edge cases, that the test suite itself might not cover.

### Do parallel coding agents need their own sandboxes?

Yes. Agents sharing a single local checkout will overwrite or mix each other’s changes. Isolated sandboxes, ideally booted from a common pre-configured snapshot, keep each agent’s work separate until it’s reviewed and ready to integrate.

### Why do merge conflicts happen even with isolated sandboxes?

Isolation prevents agents from interfering with each other in real time, but multiple agents starting from the same base snapshot often edit the same shared files, like a common test file, independently. Those conflicts only surface when a separate integration stage tries to merge everything into one branch.
