# Logbook of a Spec-Driven Developer

> Source: <https://dev.to/sidiar/logbook-of-a-spec-driven-developer-pcm>
> Published: 2026-09-30 07:32:22+00:00

*Building a 45k-line app with AI agents: what worked, and what broke*

*“Now that we have AI, why can’t we do this project in 3 weeks instead of 3 months?”*

I had a very good year with AI, achieving things in areas that were new to me much faster than the traditional way. Suddenly everybody was excited: delivery to production had multiplied, and things postponed for years, like technical debt, were finally getting out of the way. It didn’t take long for product managers and executives to notice, and to start asking questions like the one above.

As a senior engineer, I am very clear that I am 100% responsible for the code that reaches production. When someone from product raises their eyebrows and asks to do things very fast, my legs shake. The AI can spit out a mountain of code without breaking a sweat, but it is not the AI who assumes the consequences of a work badly done, of an application that does not align with the business goals, or of a random code without cohesion between its parts.

So we are struggling between 2 dimensions: Speed & Quality. How effective can we be before quality starts to decrease? What I mean by quality is alignment with business goals and alignment with architecture vision. This is where I thought Spec Driven Development (SDD) could solve both problems: it lets you **build fast** what the team has decided to and **how** the team has decided. Before answering that question at work, I wanted to find out on my own. So I proposed myself this experience in solitary, to see if that promise holds.

First, one line about me: I have been building software for 25 years, the last ones as senior engineer and engineering manager. For this experiment I used the BMad method, one of the SDD frameworks: you produce a Project Brief, then the Product Requirements, then the Architecture, and only then the AI implements, story by story. The stack is React and TypeScript with a typed-array simulation engine, and the result is live at [game-of-life-studio.com](https://game-of-life-studio.com?utm_source=devto).

Why this project? I had run a workshop a year before at my children’s school related to [Conway’s game of life](https://en.wikipedia.org/wiki/Conway%27s_Game_of_Life). Something to share and have fun with my kids, free user testing, and who knows what this can bring in the future out there. As time passed during the project I was happy to see that it had the right level of complexity needed to prove that the quality expected was not trivial. It was also fun to see how my kids were asking every day, “Is there anything new that we can see?”. They were my Product Managers!

Shaping a perfectly specified product before starting to build it has always been a big challenge and has required experienced cross functional team roles and responsibilities. Working with SDD methodology helped me a lot in sorting this out. BMad skills forced me to cover each detail of the product requirements, beginning with the Project Brief, following with Product Requirements and lastly with the Architecture. Each phase had a clear Acceptance criteria that could not be skipped until considered done by the Framework. Thanks to this, very few gaps could have sneaked into the Execution phase. AI is incredibly efficient on detecting missing requirements, edge cases and contradictions in a spec.

But this precision has a price. It became hard to wrap up the PRD phase after having the Architecture advanced, the issues kept coming out, regressions, etc. I felt stuck and unable to move on, since it was mandatory to resolve any issue, even the ones marked as low priority. I believe that we typically defer some decisions that are not so relevant, and this was not allowed — it sure is something to improve with SDD. I also had to stop several times to build and optimize a project-context file, since it started to consume enormous amounts of tokens as everything grew.

The other price was communication. The AI is like an employee that is eager for a promotion: it outputs too much content, full of references like AR-33, AR-32 or UX-DR20 in the same paragraph, and poetic terms — it took me weeks to read “Roster” without thinking of an Alice in Chains song. That story deserves its own article.

I remember thinking “I can’t wait to start the execution phase, and see how fast I can get the application implemented with best practices.”

Things didn’t turn out too bright at the beginning. In Epic 1 the process was very slow and got me absorbed in front of the keyboard, after all that long planning phase. On one side because of the communication style of the AI, and on the other side because some parts of the new code purpose could not be understood until implementation got further. At some point I found myself blindly approving decisions, and I kept asking myself: what would I do in a real working environment? I am also too proud to let AI do all the coding without my intervention, even if it was following my solution design.

It was an interesting BMad proposal to create the Story specs right before the implementation and after the previous story was done, so that any deferred work can be attended and specs are up to date with it. Also, every part of this process (Create Story > Implement > Review) had to be done in a fresh session, with the review on a different model than the implementation, which made a lot of sense.

After each PR was ready I was doing a thorough review and asking for changes and improvements, although I have to say that from the beginning I was really happy with the resulting code. Then I realized that it was me slowing down the process. The whole idea of SDD is that, on execution phase, this should go very fast.

My first decision before starting epic 2 was to split the review into two moments. During the epic, I would do only a quick review of each PR before merging it, without requesting changes (at most taking notes), so I wouldn’t slow down the implementation of the whole epic. Once the epic was done, I would take the time for a deeper review: create a refactoring branch and introduce improvements through vibe coding. Not because I did not trust the code, it looked very good, but this became a personal challenge and at the same time a good opportunity to keep up with the details of the implementation. In order to keep track of what were my contributions I would prefix each commit with “HITL refactor:” (HITL = Human In The Loop), and the PR was not merged with Squash, since I wanted to keep these improvements individually. Once I was happy, or I just did not want to keep dedicating time to this, I would create a PR that would be reviewed by Opus. Lovely!

My second decision was to accelerate the process. I had studied the Ralph loop (“let it run overnight and see in the morning”) and that was my original intention for epic 2. But I feel like this is not right, unless you are working in an experimental repository or a PoC. If you are working on important code that will be shipped to production, you should always review story by story and approve. Remember: it is not the AI who assumes the consequences.

So instead I created the “implement next story” skill. What the skill does: it takes the next story of a sprint board and turns it into a pull request, spawning three fresh subagents in turn — create the story, implement it, review it on a different model — and stops with the PR open for one human to merge. It implements nothing itself; it orchestrates, reads artifacts off disk, and refuses to guess.

This was an inflexion point for the project. This is when I finally started to move fast and deliver good work at the same time, what we all want to achieve with AI.

At the middle of epic 2 I had the idea to make the skill gather execution stats, taken from the runtime’s own transcripts — measured, not estimated. Thirty stories went through the pipeline. The median from sprint board to open PR was 67 minutes (range 43–106 across the uninterrupted runs): 10 minutes to create the story, 26 to implement it, 32 to review it and open the PR. The median story produced 180k output tokens — and 82.6 million cache-read tokens, 97% of the total, which is what makes the bill reasonable. Thirteen stories were implemented on Sonnet and reviewed by Opus; seventeen on Opus, reviewed by Fable after it came out (the first seven had to settle for a Sonnet review, and I will come back to why that matters). Never the same model twice.

Things started to move fast and smooth as I got to Epic 3, but there was still a lot of work ahead. I was lucky that Epic 3 and Epic 4 did not depend on each other mostly, because I hadn’t considered parallelization during the planning (my bad!). So I adapted the skill in order to run 2 agents at the same time, one per each epic. I learned about git worktree and iteratively improved the process, dealing with some issues on the way.

Not everything went smooth, and I think this is the most useful part of this article. Almost every rule the skill has today is there because something went wrong first.

**Two agents, one folder.** The first day I ran two epics in parallel, I launched both agents from the same folder. Each one checked at the start that there were no pending changes, and there weren’t, because the other one hadn’t written anything yet. A few minutes later the agent working on story 3.8 made its commit, and it included the status lines that the other agent had just written for story 4.1. One story’s commit carried changes that belonged to the other. That’s how I learned about git worktree: one folder per lane. The lesson I take from it applies anywhere: a check that runs at the start only protects you from the state at the start. If something must be true, check it again right before any step you can’t undo, like branching or committing.

A few days later I did it again (my bad!). I opened two terminals in the main folder, even though one of the epics already had its own worktree. Which epic lived in which folder was only in my head. Now the skill figures that out by itself, and it locks the folder so a second session stops before writing anything.

**The reviewer reviewing itself.** The whole point of the review step is that a different model looks at the code. The first version of the skill always used Opus as reviewer, so when Opus implemented a story, Opus reviewed its own work. Nothing failed and nothing warned me; it just silently stopped being a second opinion. My first fix, “use the other model”, created a new problem. The most important stories, the ones I sent to Opus because the rest would copy their patterns, ended up reviewed by the weakest model. Now it’s a table: Sonnet is reviewed by Opus, Opus is reviewed by Fable. Before starting the review, the skill has to say which model implemented and which one will review. A silent mistake became a visible one.

**The 65-hour review.** When I added the stats table, one story showed a review phase of 65 hours. The AI didn’t spend 65 hours reviewing. The clock was counting nights, usage-limit resets, and the time the PR was waiting for me to make a decision. Now the stats count only active time and list every gap they left out. It was also the first sign of something I’ll come back to at the end: the slowest part of the pipeline was not the AI anymore.

This was for me a very good start as a Spec-Driven Developer, and I am sure this experience will help me to get the best of AI and achieve better results faster. However I worked on my own, and I am now eager to try this in a real working environment, with a team and an active business going on. I can tell you that this is not a comfortable path: setting all up for the execution takes time and patience, and I am sure you will be asked several times, “When is this going to be ready?”. One advice for each aspect of the process:

I couldn’t have ever built an app like this in this short time, and I am impressed with how the resulting code looks. But there is one irony to close with: by the end, the slowest part of the pipeline was not the AI anymore. It was me, deciding on open PRs, and the CI queue behind them. That is a very different problem from the one I started with, and I believe it is the problem many teams are about to meet.

So, to the product manager asking for 3 weeks instead of 3 months: yes, the code can be written much faster, but only after the specs are done, and not faster than the team can review and decide. That is where the time goes now.

The app is live at [game-of-life-studio.com](https://game-of-life-studio.com?utm_source=devto), the code is at [github.com/sidiar/game-of-life-studio](https://github.com/sidiar/game-of-life-studio), and the skill has its own repo at [github.com/sidiar/implement-next-story](https://github.com/sidiar/implement-next-story). If your team is meeting this problem too, I would love to hear how you are dealing with it.
