Spec-Driven Development with AI Agents: Spec Kit vs OpenSpec vs BMAD (Plus a Hands-On OpenSpec… A hands-on comparison of three spec-driven development tools for AI coding agents — GitHub Spec Kit, OpenSpec and the BMAD Method — finds they share the goal of producing a persistent, reviewable paper trail that outlives a single chat but differ sharply in weight and ceremony. The author names OpenSpec, whose lifecycle runs Explore → Propose → Review → Apply → Archive and whose changes are folders of agent-generated proposal.md, design.md, tasks.md and delta specs, as the favourite and lightest option, best suited to solo developers and small teams on existing codebases. GitHub Spec Kit is described as the heaviest of the lightweight options, with a constitution, specification, plan, tasks and implement pipeline, while BMAD simulates an entire agile team of analyst, product manager, architect, scrum master, developer and QA personas and is best for larger or greenfield projects. Hey there Let’s talk about something that I tried my hands on recently and has changed how I work with AI coding agents. AI coding agents are brilliant at writing code. They’re much less brilliant at remembering why they wrote it. Ask one to build a feature in a single chat and you’ll often get something that looks great but quietly misses half your requirements. And once the chat is gone, so is all the reasoning behind it. Sound familiar? A new family of tools has popped up to fix exactly this. They put a structured workflow between your idea and the code. The three I keep seeing are GitHub Spec Kit , OpenSpec https://github.com/Fission-AI/OpenSpec and the BMAD Method https://github.com/bmad-code-org/BMAD-METHOD . They share a goal but feel very different in weight, ceremony and philosophy. In this post I’ll compare them, and then we’ll build a small first project together with the lightest one. You’ll often hear these tools called “spec-driven development”. But in practice, I’ve found the spec isn’t really the starting point . It’s more like an output of the workflow. You start with a conversation. The agent then writes the specs for you, so there’s no hand-writing PRDs Product Requirements Documents, the documents that describe what a product should do and why . What you get is a persistent, reviewable paper trail that outlives any single chat. That’s the real magic. Spec Kit is GitHub’s own toolkit, with a CLI and a set of slash commands. The pipeline is clear: set your project principles the “constitution” , write a specification, plan the technical approach, break it into tasks, then implement. There are optional steps for clearing up ambiguity and checking consistency too. What’s great: It’s well structured, backed by GitHub, and works with lots of agents. The constitution gives you project-wide rules that every change has to respect. What to watch for: It’s the heaviest of the lightweight options, with quite a bit of upfront ceremony. For small changes it can feel like overkill. And a small confession: the word “constitution” just grated on me. Trivial, I know, but it’s real 😄 Best for: Teams who want formal governance and consistency across many contributors. OpenSpec is the minimalist, and honestly my favourite of the bunch. Its lifecycle is simple: Explore → Propose → Review → Apply → Archive . Each piece of work is called a “change”, which is just a folder of agent-generated files: proposal.md, design.md, tasks.md, plus “delta specs” describing what’s been added, modified or removed. When you archive a change, those deltas fold into your main specs, so they always reflect the system as it’s actually built . What’s great: What to watch for: It’s less opinionated about team process, and your specs can drift if you skip the sync/archive step. Best for: Solo developers/scientist like me and small teams working on existing codebases. BMAD Breakthrough Method for Agile AI-Driven Development takes a completely different route. Instead of a light spec loop, it simulates an entire agile team. You get specialised agent personas, like an analyst, product manager, architect, scrum master, developer and QA. Each one produces the artifacts for their role: a PRD, an architecture document, then detailed stories for the developer agent to implement. What’s great: It’s comprehensive, covering everything from ideation to delivery, and it shines on greenfield projects where you want rigorous planning and clear handoffs. What to watch for: It’s the heaviest by a distance. There’s lots to learn and lots of documents to review, which is overkill for a small feature or a quick fix. Best for: Larger or greenfield projects, and anyone who loves thinking in agile roles. My own rule of thumb is to match the tool to the size of the work. A full agile simulation for a one-day feature is wasted effort, and a thin workflow for a multi-month product might leave gaps. One quick note: these tools evolve fast, so double-check each project’s docs for the latest commands and features. Ready to get hands-on? I will show you building a very tiny and simple Scientific Paper Tracker that I find very handy. It’s small enough that you can focus on learning the workflow instead of fighting the app. Also available at my GitHub Step 1: Check Node.js In VS Code create a new folder and open the bash terminal. Check Node.js version. You’ll need Node.js 20.19 or newer. Step 2: Install OpenSpec Step 3: Create your project Step 4: Initialise OpenSpec Follow the pop-up instruction on terminal and select your coding agent Claude Code, Codex, Cursor, etc. . I have used Codex for this demo. Step 5: Open the project OpenSpec will create an openspec/ folder and the integration files for your agent. You should see roughly this: 6. Explore the idea Open Coding Agent-CLI for me its Codex . In your agent’s chat, tell it what you’re after: Choose the LLM-model while typing /model in Codex-CLI . I selected GPT-6.1-Sol-medium which I find judicious choice in terms of speed and reasoning. Also check the $openspec-explore which provides lots of options for pre-prepared documentation. I am not covering them all, just try it out and check what works best for your requitements. Remember, explore is thinking only . No code gets written yet, so just enjoy the conversation. Step 7. Create the proposal For instance I provided my agent following propsal Once the idea feels solid, say: You’ll typically end up with: Step 8. Review before coding This is my favourite step. Read through proposal.md, specs/ and tasks.md, and check that the requirements are simple and match what you want. I am sharing them all here for reproducibility. Proposal WhyUsers need a simple way to remember scientific papers and track which ones they have read. A small, user-friendly web page can provide this without accounts or a complex research-management system. What Changes- Add a single web page with a form for a paper's title, authors, and topic.- Display saved papers with their metadata and a clear read or unread status.- Provide actions to mark a paper as read and delete a paper.- Keep papers across page reloads using storage in the current browser.- Provide a helpful empty state, accessible controls, and a layout that works on phones and desktops.- Keep the first version limited to these actions; accounts, synchronization, search, editing, PDF uploads, and metadata lookup are outside scope. Capabilities New Capabilities- paper-library : Add, list, retain, mark as read, and delete scientific paper records through a simple web interface. Modified CapabilitiesNone. ImpactThe project currently contains OpenSpec setup only. This change introduces a static web page using HTML, CSS, and JavaScript, with browser local storage and no runtime dependencies, backend, database service, or external API. Records belong to the current browser and origin; clearing browser storage removes them. Tasks 1. Static page and layout- x 1.1 Create index.html, styles.css, and app.js with a heading, browser-storage note, visibly labeled title/authors/topic form, status area, and empty paper list; verify the page loads through a static server without missing assets.- 1.2 Style the form and list with readable spacing, visible keyboard focus, wrapping metadata, and controls that stack on narrow screens; verify keyboard navigation and no horizontal page scrolling at 320 CSS pixels.- x 1.3 Add README.md with a simple static-server command and browser-local storage limitations; verify the documented command serves the page at the documented address. 2. Paper records and persistence- x 2.1 Load and validate the saved JSON array using the paper-tracker.papers localStorage key; verify a missing key shows the empty state and unreadable or malformed data shows an error without overwriting data or enabling mutations.- x 2.2 Implement add and list behavior with unique record IDs, trimmed fields, required nonblank title, optional authors/topic, unread default, insertion order, and literal text rendering; verify complete and title-only submissions, whitespace rejection, repeated titles, form reset after success, and harmless HTML-like text.- x 2.3 Save proposed mutations before updating displayed state; verify reload retains added records and a failed write leaves the previous list and form contents intact with a visible error.- x 2.4 Add the add/list/persistence checks to a short manual verification checklist in README.md; verify the checklist includes invalid stored data and failed writes without adding a test framework. 3. Reading and deletion actions- x 3.1 Add a mark-as-read action for unread records and display a textual Read status; verify only the selected record changes, metadata stays intact, and the status survives reload.- x 3.2 Add immediate deletion by record ID; verify one of two equal-title records can be removed independently, deletion persists after reload, and deleting the final record restores the empty state.- 3.3 Announce action results and errors and restore useful focus when an action removes the focused control; verify both actions work with a keyboard and focus remains usable.- x 3.4 Extend README.md usage and verification notes for reading and deletion; verify the steps describe actual button labels and the lack of undo. 4. Integrated browser verification- 4.1 Run the full checklist in a desktop browser and at a 320 CSS pixel viewport with long metadata: add repeated titles, mark one read, delete another, reload, and delete the last; verify the saved state, focus behavior, messages, and layout together and record the browser and result. Design ContextThis is a greenfield project: only OpenSpec configuration and workflow skills exist, with no application, tests, or existing capability specs. See proposal.md for motivation and specs/paper-library/spec.md for behavior. A short design is useful to establish the first application structure and storage decisions before coding. Goals / Non-Goals Goals: Use a small static application, native browser controls, and a single persistence mechanism. Keep setup and maintenance minimal. Non-Goals: Introduce a framework, build pipeline, server, dependency management, routing, or multi-user data model. Decisions Plain static filesUse root-level index.html, styles.css, and app.js. A locally served static page has a stable origin for browser storage and can later be hosted without changing the architecture. Add a short README with a simple static-server command and usage instructions. Prefer these files over a frontend framework because the page has one form and one list, with no routing or complex state. Browser-local JSON storageStore an array under one localStorage key, paper-tracker.papers. Each record contains a generated unique id, title, authors, topic, and a read boolean. Authors and topic are plain strings; do not introduce author entities or topic taxonomies. Preserve insertion order. Use IDs for actions so repeated titles are safe.Read storage once on startup. A missing key means an empty library. Handle unreadable or malformed data with a visible error and disable mutations to avoid overwriting existing data; do not silently clear it. For mutations, build the next array and save it before updating the displayed state. If saving fails, retain the old list and form contents and report the failure. A server or database would add operations and synchronization concerns beyond this version. One form and a simple listPlace a heading, a short browser-storage note, the add form, and the paper list on one page. Require only the title, explicitly label the other fields optional, and clear the form after a successful save. Use readable spacing, a restrained color palette, and native buttons with visible focus. Each paper shows metadata, a text status, a mark-as-read button when unread, and a delete button. Omit empty optional metadata. Render user text with textContent. Wrap long values and stack controls on narrow screens.Deletion takes effect immediately with no extra dialog, keeping interaction short. Marking a read paper unread and undoing deletion are outside this version. After an action removes a focused control, move focus to a nearby available control or the title field. Announce action results and errors through a small accessible status area. Proportional verificationUse a browser checklist covering form validation, repeated titles, independent actions, persistence, empty state, literal text rendering, keyboard use, and a narrow viewport. Also check unavailable storage and malformed stored data. Avoid adding a test framework or build tools solely for this small first version. Risks / Trade-offs- Browser storage can be cleared and is isolated by browser and origin - Explain this on the page and in README; no synchronization or backup is promised.- Storage can fail or contain malformed data - Surface errors, preserve saved data, and do not claim unsuccessful mutations were saved.- Immediate deletion has no undo - Use a clearly labeled Delete button; accept this limitation for the minimal first version.- Other tabs can write the same storage - Treat this as a single-tab tool; concurrent-tab synchronization is outside scope. Migration PlanThere is no existing application or data to migrate. Serve the static files locally for verification; production hosting is outside this change. Roll back by removing the new application files. Retain browser data unless the user explicitly chooses to clear it. Reviewing before any code exists is one of the most valuable habits in this whole workflow. If something’s off, edit it yourself or ask the agent to do. Step 9. Apply the change Start a fresh chat everything the agent needs is saved on disk and say: Then sit back and watch the agent work through tasks.md, ticking off each item as it goes. It’s surprisingly satisfying Step 10: Test and archive Once everything works: OpenSpec moves the change into openspec/changes/archive/ and updates openspec/specs/ so your specs describe the system exactly as it was built. Pro tip that I learned after messing it up in my first try : do this before merging your PR. Otherwise the sync and archive changes will need a separate PR. The whole flow basically summarise as follow: IDEA → Explore → Propose → proposal.md, specs/, design.md, tasks.md → Review → Apply → CODE → Test → Archive → openspec/specs/ A suggested stack for the demo If you’re not sure what to build it with, this combo keeps things simple: OpenSpec + Python + FastAPI + SQLite With just four endpoints: POST /papersGET /papersPATCH /papers/{id}DELETE /papers/{id} A quick note on commands: the exact slash-command spelling depends on your coding agent. openspec init installs the right commands or skills for your tool. Claude Code commonly shows /opsx:…, while other tools expose OpenSpec through skills. And now the reveal of my Scientific Paper Tracker, which looks like this: Next steps Once your tracker works, try a second change, like adding tags or search. Modify it the way you would like your app to behave. Point the agent at the archived spec from your first change and notice how quickly it picks up the context. That “specs as context primers” effect is where OpenSpec really starts to pay off. If you remember only a few things from this post, make it these: AI coding agents write code fast, but without structure they forget why they wrote it. Spec-driven tools fix that by putting a persistent, reviewable workflow between your idea and the code. The big lesson is that the specs are a by-product of the workflow, not the starting point . You start with a conversation, the agent writes the artifacts, and you review them before any code exists. That review step, plus archiving each finished change so your specs always match what was really built, is what makes this approach pay off. My practical advice: match the tool to the size of the work and use the lightest workflow that gives you enough structure. If you’re new to all this, try OpenSpec on a small project like the Paper Tracker. It takes under an hour, and by your second change you’ll see how much faster your agent picks up context. I’d genuinely love to hear from you. After reading, drop a constructive feedback/ comment below and tell me: Your experiences will help other readers pick the right approach, and I read every single comment. If you enjoyed this post, a clap 👏 and a share with a colleague who’s curious about AI-driven development would mean the world and also keep me motivated. Thanks for reading, and happy building Spec-Driven Development with AI Agents: Spec Kit vs OpenSpec vs BMAD Plus a Hands-On OpenSpec… https://pub.towardsai.net/spec-driven-development-with-ai-agents-spec-kit-vs-openspec-vs-bmad-plus-a-hands-on-openspec-7b6997074582 was originally published in Towards AI https://pub.towardsai.net on Medium, where people are continuing the conversation by highlighting and responding to this story.