Production-ready software development at the speed of thought After six years of research, four years working with LLMs, and almost a year of agentic coding in a production-ready environment, software developer and researcher Osequi concluded that LLMs with the right guardrails and tools can safely automate the design, implementation, and verification of production-ready software, leaving humans to understand the problem and steer the machine toward a solution. In January 2026, after the Claude Code hype settled, Osequi began implementing a production-ready version of the Olog editor, aiming to "vibecode" the new web app with strong guardrails using battle-tested React, Next.js and TypeScript templates. The study covers the implementation and verification phases of the software development life cycle, arguing that almost every SDLC step is formally or semi-formally specifiable and thus verifiable cost-effectively by new tools and services. Speed of thought LLMs—with the right guardrails and tools—can safely automate the design, implementation, and verification of production-ready software. What’s left for humans is to understand the problem and steer the machine toward a solution. This is production-ready software development at the speed of thought. Background After six years of research, four years working with LLMs, almost a year of agentic coding in a production-ready environment, I found that: - LLMs are truly useful for generating verifiable outputs - Almost every step in the SDLC process is formally or semi-formally specifiable, and thus verifiable in a cost-effective way by new tools and services LLMs do the heavy lifting. They are indispensable when creating production-ready software. Research In autumn 2022, when ChatGPT came out, I started a journey to find a method for creating likely-correct software https://www.osequi.com/slides/likely-correct-software/likely-correct-software.html through rapid iterations https://www.osequi.com/slides/rapid-iteration/rapid-iteration.html . Last year, I presented a study https://www.osequi.com/studies/list/list.html on how to use a mix of formal and semi-formal methods to achieve likely-correct understanding and design—the first two phases of the software development life cycle SDLC . This year I cover the next two phases: implementation and verification—again—the likely-correct way. The remaining phases, deployment and maintenance, are not relevant in this context. The goal of these studies is to find out what is possible in terms of correctness, and at what cost, in software development. To find a pragmatic approach to LLMs, to discover where they are truly useful and what we can do to make them truly useful. To eliminate the unpleasant surprises they often throw at us. The final goal is to create better software faster. Production-ready code, faster There is no exact definition of what production-ready software is or how to create it. There is broad agreement that such software is: - Fully tested—from specifications to code—throughout all the phases of the SDLC - It is built for the long run. It’s not the first iteration. That would be an MVP, a prototype, a proof of concept - It is designed to evolve through rapid iterations and to face the challenges of real-world, large-scale use You can see an example of production-ready code in A likely-correct list https://www.osequi.com/studies/list/list.html , a prequel to this current study. The study presents: - The main aspects Information Architecture , the process and the deliverables SDLC of a production-ready code - How these parts relate to formalism, correctness and cost - How, and which, formal and semi-formal methods ensure likely-correct deliverables and rapid iterations The study concludes: And in a future stage, it should provide AI/ML practitioners with insights to combine generative AI’s speed with the rigor of these new techniques to produce better software, faster. Now let’s pick up from here. In January 2026, after the Claude Code hype settled down, I decided to start the implementation of the production-ready version of the Olog editor—a project where the MVP received positive feedback and in which I saw potential for further investment. In my workflow, ologs ensure a likely-correct understanding of a problem domain. The idea was to vibecode the new web app with strong guardrails. I had on hand a strong set of React, Next.js and Typescript templates created in Silicon Valley environments and battle-tested in production, ensuring that the written code was of high quality. I had some strong opinions about an ideal code architecture. Organizing and maintaining a codebase is a recurring top pain point https://2023.stateofjs.com/en-US/usage/ top js pain points for front-end developers and I had spent endless effort http://metamn.io/react/ figuring it out. If it’s a pain point for humans it will be a pain point for LLMs. I asked the LLMs for help, and with three American and two Chinese models, we came up with such a solid production-ready architecture that they said they had never seen something similar before. Lol 😀 In the end, we’ve put together a generic multi-dimensional guardrail system in which: - The columns represent the app features - The rows represent the code architecture layers - A third, controlling dimension is error handling With this toolset—my best knowledge ever—I started vibecoding the features one by one. A detailed workflow document would instruct the LLM how to implement a feature step by step across the architectural layers. My job was easy: - Feed in the specs - Make sure all the steps in the workflow are fully executed - Run all tests and achieve close to 100% code coverage - Verify the results live then create the end-to-end tests We—the junior dev team provided by the LLM and I—were flying high. The speed was right. The codebase looked promising. The investment in the code architecture and the templates paid off. Even Simon Willison praised https://news.ycombinator.com/item?id=48524489 the results. On the other hand, the product experience was strange. Subtle user interface errors, messy business logic, once-fixed problems resurfacing, code smells, inconsistency, duplicated code, ghost code. I had a bad feeling— We were not ready for production https://news.ycombinator.com/item?id=48421559 . Also the Claude experience turned out to be awful. As the problems within the codebase grew that intellectual, tortured genius became more stubborn, over-refusing and always lecturing. What’s next? The problem was that while the primary focus—implementing the specs—went almost flawlessly, the secondary focus—following the architectural layers and coding rules—drifted at a subtle level while it was looking good on the surface. To fix this we’ve started a second iteration now focusing on the rows of the guardrail matrix, on the non-user-facing coding and organizational rules that make a product robust. We also chose a much better harness Opencode and a coding agent with a dedicated engineering mindset Deepseek Flash at a fraction of the price and hassle. Also changed the rules of the game with LLMs: Instruct — Never ask — Always verify. - From this point on, throughout the SDLC process, I’ll never ask for advice or brainstorm. I’ll handle this part myself using the old methods. - From now on, I give LLMs plain instructions to produce bite-sized chunks of code that I can skim for errors ghost code, duplicated code, etc without too much effort. The result is more than promising. After two perpendicular iterations and an overarching, integrating error-management implementation I’ve got production-ready code: it follows the specs, it follows the rules, smells good, looks good, it is fully tested and easy to iterate on. I can affirm that the implementation and verification SDLC phases are safely doable with LLMs when the right guardrails, coding agents, and human supervisors are in place. Why LLMs for I+V? - They write code and tests as well as humans, but orders of magnitude faster - They take these daunting, tiresome tasks completely off our shoulders - Writing code manually is now obsolete just as writing assembly code became obsolete when C, BASIC, and Pascal came onto the scene - Now software engineers can focus on higher levels of abstraction: pure problem solving vs. tiresome, good-old plumbing Better production-ready code, faster While I was figuring out the magic formula ... SDLC + LLM = U + LikelyCorrect D + SafelyAutomated I + V ... others reached the same or even better conclusions and created useful tools. The age-old practice of Contract programming Design by contract https://en.wikipedia.org/wiki/Design by contract took off in a new form—Vericode—and enhanced BDD to produce formally verified code https://scidonia.ai/blog/from-bdd-to-proof/ . A React/Typescript implementation https://midspiral.com/ , which I’ll use in my next projects, comes with this thesis: 1. Humans define what should be built 2. AI handles implementation 3. Machines guarantee correctness Meanwhile, Shopify went even further https://shopify.engineering/shop-app-migration . They’ve managed to get the agents write specs although from an existing production-ready app and its codebase and implementation plans; then execute and verify these plans. Before you get too excited https://www.youtube.com/watch?v=vDjW dRyKXY I should fill in the details about the invisible and hard work behind the scenes https://shopify.engineering/helix that makes that possible. According to Shopify: - Getting consistent, high-quality, and maintainable results out of the LLM black box is difficult - While the generated code satisfies the feature requirements it still introduces duplication, architectural drift, and performance problems - Highly opinionated architecture, test coverage, rapid iterations verifiable chunks of output , visual and adversarial reviews are the guardrails that make this picture wondrous . Production-ready software at the speed of thought In March 2024, the CIA presented a report https://www.cia.gov/resources/csi/studies-in-intelligence/studies-in-intelligence-68-no-1-extracts-march-2024/future-of-intelligence-the-incalculable-element-the-promise-and-peril-of-artificial-intelligence/ on the promise and peril of artificial intelligence concluding “Generative AI is neither quite so wondrous nor quite so bleak”. In autumn 2026, I can confirm this. You cannot change the world with a single prompt, but you can safely automate most of the software development process. Today in the SDLC+LLM process D+I+V are— formally and semi-formally —solved. Now agentic software development reduces to these steps: 1. You start by understanding the problem domain 2. AI helps you by sketching out the ontology and taxonomy in an Olog editor 3. In the Stately FSM editor https://stately.ai/ , you model together how these parts behave and interact 4. When the solution to the problem is clear you design the user interface and experience by using the concept design DSL 5. Finally you let it loose: Using the guardrails and the vericoding techniques the AI creates the production-ready code This is software development at the speed of thought. Resources 1. Likely correct software https://www.osequi.com/slides/likely-correct-software/likely-correct-software.html / — Osequi, 2023 2. Rapid iteration in software development https://www.osequi.com/slides/rapid-iteration/rapid-iteration.html / — Osequi, 2023 3. A likely-correct list https://www.osequi.com/studies/list/list.html — Osequi, 2025 4. Future of Intelligence: “The Incalculable Element”: The Promise and Peril of Artificial Intelligence https://www.cia.gov/resources/csi/studies-in-intelligence/studies-in-intelligence-68-no-1-extracts-march-2024/future-of-intelligence-the-incalculable-element-the-promise-and-peril-of-artificial-intelligence/ — CIA, 2024 5. JavaScript Pain Points https://2023.stateofjs.com/en-US/usage/ top js pain points — Stack Overflow, State Of Javascript, 2023 6. To React with best practices http://metamn.io/react/ — Metamn, 2018-2021 7. A Hacker News comment thread on Software Architecture Guide https://news.ycombinator.com/item?id=48524489 — Simon Willison, 2026 8. A Hacker News comment thread on: Ask HN: Why is the HN crowd so anti-AI? https://news.ycombinator.com/item?id=48421559 — The author, 2026 9. Design by contract https://en.wikipedia.org/wiki/Design by contract — Wikipedia 10. From Vibecoding to Vericoding: A Gradient, Not a Jump https://scidonia.ai/blog/from-bdd-to-proof/ — Scidonia, 2026 11. Midspiral https://midspiral.com/ — 2026 12. Migrating Shop app from React Native to native https://shopify.engineering/shop-app-migration — Shopify, September 2026 13. Rails World 2026 Opening Keynote https://www.youtube.com/watch?v=vDjW dRyKXY — DHH, September 2026 14. Helix: The internal tool powering our Shopify app's native migration https://shopify.engineering/helix — Shopify, September 2026 15. Stately https://stately.ai/ — 2026 About the author The author http://metamn.io/ holds a degree in mathematics and computer science. He is a self-taught UX/UI designer with works featured in online galleries. Recently, he has been running a research and development studio https://www.osequi.com/ specializing in software correctness and rapid iteration, providing consulting services to companies. Credits This document was created using Notion and published using a modified version of Tufte CSS. It was checked with the Google Docs spell checker and the write-good naive linter for English prose. Then LLMs corrected the non-native English words and constructions—I’m inherently prone to using—to your delight.