cd /news/artificial-intelligence/production-ready-software-developmen… · home › topics › artificial-intelligence › article
[ARTICLE · art-143094] src=osequi.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Production-ready software development at the speed of thought

After six years of research, four years working with LLMs, and almost a year of agentic coding in a production-ready environment, software developer and researcher Osequi concluded that LLMs with the right guardrails and tools can safely automate the design, implementation, and verification of production-ready software, leaving humans to understand the problem and steer the machine toward a solution. In January 2026, after the Claude Code hype settled, Osequi began implementing a production-ready version of the Olog editor, aiming to "vibecode" the new web app with strong guardrails using battle-tested React, Next.js and TypeScript templates. The study covers the implementation and verification phases of the software development life cycle, arguing that almost every SDLC step is formally or semi-formally specifiable and thus verifiable cost-effectively by new tools and services.

read9 min views1 publishedOct 1, 2026

LLMs—with the right guardrails and tools—can safely automate the design, implementation, and verification of production-ready software. What’s left for humans is to understand the problem and steer the machine toward a solution. This is production-ready software development at the speed of thought.

Background #

After six years of research, four years working with LLMs, almost a year of agentic coding in a production-ready environment, I found that:

  • LLMs are truly useful for generating verifiable outputs
  • Almost every step in the SDLC process is formally or semi-formally specifiable, and thus verifiable in a cost-effective way by new tools and services

LLMs do the heavy lifting. They are indispensable when creating production-ready software.

Research #

      In autumn 2022, when ChatGPT came out, I started a journey to find a
      method for creating
          [likely-correct software](https://www.osequi.com/slides/likely-correct-software/likely-correct-software.html)
          through
          [rapid iterations](https://www.osequi.com/slides/rapid-iteration/rapid-iteration.html).
        

          Last year, I presented a
          [study](https://www.osequi.com/studies/list/list.html) on
          how to use a mix of formal and semi-formal methods to achieve
          likely-correct understanding and design—the first two phases of the
          software development life cycle (SDLC).

This year I cover the next two phases: implementation and verification—again—the likely-correct way. The remaining phases, deployment and maintenance, are not relevant in this context.

The goal of these studies is to find out what is possible in terms of correctness, and at what cost, in software development. To find a pragmatic approach to LLMs, to discover where they are truly useful and what we can do to make them truly useful. To eliminate the unpleasant surprises they often throw at us.

The final goal is to create better software faster.

Production-ready code, faster #

There is no exact definition of what production-ready software is or how to create it. There is broad agreement that such software is:

  • Fully tested—from specifications to code—throughout all the phases of the SDLC

  • It is built for the long run. It’s not the first iteration. That would be an MVP, a prototype, a proof of concept

  • It is designed to evolve through rapid iterations and to face the challenges of real-world, large-scale use

        You can see an example of production-ready code in
    
      
        [A likely-correct list](https://www.osequi.com/studies/list/list.html), a prequel to this current study.
    

The study presents:

  • The main aspects (Information Architecture), the process and the deliverables (SDLC) of a production-ready code
  • How these parts relate to formalism, correctness and cost
  • How, and which, formal and semi-formal methods ensure likely-correct deliverables and rapid iterations

The study concludes:

        And in a future stage, it should provide AI/ML practitioners with
        insights to combine generative AI’s speed with the rigor of these

      
        new techniques to produce *better* software, faster.

Now let’s pick up from here.

In January 2026, after the Claude Code hype settled down, I decided to start the implementation of the production-ready version of the Olog editor—a project where the MVP received positive feedback and in which I saw potential for further investment. In my workflow, ologs ensure a likely-correct understanding of a problem domain.

The idea was to vibecode the new web app with strong guardrails.

I had on hand a strong set of React, Next.js and Typescript templates created in Silicon Valley environments and battle-tested in production, ensuring that the written code was of high quality.

      I had some strong opinions about an ideal code architecture.
      Organizing and maintaining a codebase is
          [a recurring top pain point](https://2023.stateofjs.com/en-US/usage/#top_js_pain_points)
          for front-end developers and I had spent
          [endless effort](http://metamn.io/react/) figuring it out.
        

If it’s a pain point for humans it will be a pain point for LLMs. I asked the LLMs for help, and with three American and two Chinese models, we came up with such a solid production-ready architecture that they said they had never seen something similar before. Lol 😀

In the end, we’ve put together a generic multi-dimensional guardrail system in which:

  • The columns represent the app features
  • The rows represent the code architecture layers
  • A third, controlling dimension is error handling

With this toolset—my best knowledge ever—I started vibecoding the features one by one.

A detailed workflow document would instruct the LLM how to implement a feature step by step across the architectural layers. My job was easy:

  • Feed in the specs
  • Make sure all the steps in the workflow are fully executed
  • Run all tests and achieve close to 100% code coverage
- Verify the results live then create the end-to-end tests

          We—the junior dev team provided by the LLM and I—were flying high. The
          speed was right. The codebase looked promising. The investment in the
          code architecture and the templates paid off. Even
          [Simon Willison praised](https://news.ycombinator.com/item?id=48524489)
          the results.
        

          On the other hand, the product experience was strange. Subtle user
          interface errors, messy business logic, once-fixed problems
          resurfacing, code smells, inconsistency, duplicated code, ghost code.
          I had a bad feeling—[We were not ready for production](https://news.ycombinator.com/item?id=48421559).

Also the Claude experience turned out to be awful. As the problems within the codebase grew that intellectual, tortured genius became more stubborn, over-refusing and always lecturing.

What’s next?

The problem was that while the primary focus—implementing the specs—went almost flawlessly, the secondary focus—following the architectural layers and coding rules—drifted at a subtle level while it was looking good on the surface.

To fix this we’ve started a second iteration now focusing on the rows of the guardrail matrix, on the non-user-facing coding and organizational rules that make a product robust.

We also chose a much better harness (Opencode) and a coding agent with a dedicated engineering mindset (Deepseek Flash) at a fraction of the price and hassle.

Also changed the rules of the game with LLMs: Instruct — Never ask — Always verify.

  • From this point on, throughout the SDLC process, I’ll never ask for advice or brainstorm. I’ll handle this part myself using the old methods.
  • From now on, I give LLMs plain instructions to produce bite-sized chunks of code that I can skim for errors (ghost code, duplicated code, etc) without too much effort.

The result is more than promising.

After two perpendicular iterations and an overarching, integrating error-management implementation I’ve got production-ready code: it follows the specs, it follows the rules, smells good, looks good, it is fully tested and easy to iterate on.

I can affirm that the implementation and verification SDLC phases are safely doable with LLMs when the right guardrails, coding agents, and human supervisors are in place.

Why LLMs for I+V?

  • They write code and tests as well as humans, but orders of magnitude faster
  • They take these daunting, tiresome tasks completely off our shoulders
  • Writing code manually is now obsolete just as writing assembly code became obsolete when C, BASIC, and Pascal came onto the scene
  • Now software engineers can focus on higher levels of abstraction: pure problem solving vs. tiresome, good-old plumbing

Better production-ready code, faster #

While I was figuring out the magic formula ...

          SDLC + LLM = U + LikelyCorrect(D) + SafelyAutomated(I + V)

... others reached the same or even better conclusions and created useful tools.

      The age-old practice of
          [Contract programming (Design by contract)](https://en.wikipedia.org/wiki/Design_by_contract)
          took off in a new form—Vericode—and enhanced
          [BDD to produce formally verified code](https://scidonia.ai/blog/from-bdd-to-proof/).
        

          A
          [React/Typescript implementation](https://midspiral.com/),
          which I’ll use in my next projects, comes with this thesis:
  1. Humans define what should be built
  2. AI handles implementation
  3. Machines guarantee correctness
          Meanwhile,
          [Shopify went even further](https://shopify.engineering/shop-app-migration).

They’ve managed to get the agents write specs (although from an existing production-ready app and its codebase) and implementation plans; then execute and verify these plans.

      Before
          [you get too excited](https://www.youtube.com/watch?v=vDjW_dRyKXY)
          I should fill in the details about the invisible and hard work
          [behind the scenes](https://shopify.engineering/helix) that
          makes that possible. According to Shopify:
  • Getting consistent, high-quality, and maintainable results out of the LLM black box is difficult
  • While the generated code satisfies the feature requirements it still introduces duplication, architectural drift, and performance problems
  •       Highly opinionated architecture, test coverage, rapid iterations
    
            (verifiable chunks of output), visual and adversarial reviews are
            the guardrails that make this picture *wondrous* .

Production-ready software at the speed of thought #

      In March 2024, the CIA presented
          [a report](https://www.cia.gov/resources/csi/studies-in-intelligence/studies-in-intelligence-68-no-1-extracts-march-2024/future-of-intelligence-the-incalculable-element-the-promise-and-peril-of-artificial-intelligence/)
          on the promise and peril of artificial intelligence concluding
          “Generative AI is neither quite so wondrous nor quite so bleak”.

In autumn 2026, I can confirm this. You cannot change the world with a single prompt, but you can safely automate most of the software development process.

      Today in the SDLC+LLM process D+I+V are—*formally* and
      *semi-formally*—solved.

Now agentic software development reduces to these steps:

  1. You start by understanding the problem domain
  2. (AI helps you) by sketching out the ontology and taxonomy in an Olog editor
            In the [Stately FSM editor](https://stately.ai/) , you
            model (together) how these parts behave and interact
  1. When the solution to the problem is clear you design the user interface and experience by using the concept design DSL
  2. Finally you let it loose: Using the guardrails and the vericoding techniques the AI creates the production-ready code

This is software development at the speed of thought.

Resources #

Likely correct software — Osequi, 2023 2. Rapid iteration in software development — Osequi, 2023 3. A likely-correct list — Osequi, 2025 4. Future of Intelligence: “The Incalculable Element”: The Promise and Peril of Artificial Intelligence — CIA, 2024 5. JavaScript Pain Points — Stack Overflow, State Of Javascript, 2023 6.

[To React with best practices](http://metamn.io/react/) —
            Metamn, 2018-2021

A Hacker News comment thread on Software Architecture Guide — Simon Willison, 2026 8. A Hacker News comment thread on: Ask HN: Why is the HN crowd so anti-AI? — The author, 2026 9. Design by contract — Wikipedia 10.

[From Vibecoding to Vericoding: A Gradient, Not a Jump](https://scidonia.ai/blog/from-bdd-to-proof/) — Scidonia, 2026
11. [Midspiral](https://midspiral.com/) — 2026

Migrating Shop app from React Native to native — Shopify, September 2026 13. Rails World 2026 Opening Keynote — DHH, September 2026 14. [Helix: The internal tool powering our Shopify app's native

              migration](https://shopify.engineering/helix) — Shopify, September 2026
15. [Stately](https://stately.ai/) — 2026

About the author #

[The author](http://metamn.io/) holds a degree in
          mathematics and computer science. He is a self-taught UX/UI designer
          with works featured in online galleries.
        

          Recently, he has been running
          [a research and development studio](https://www.osequi.com/)
          specializing in software correctness and rapid iteration, providing
          consulting services to companies.

Credits #

This document was created using Notion and published using a modified version of Tufte CSS.

      It was checked with the Google Docs spell checker and the
      `write-good` naive linter for English prose.

Then LLMs corrected the non-native English words and constructions—I’m inherently prone to using—to your delight.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @osequi 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/production-ready-sof…] indexed:0 read:9min 2026-10-01 · —