{"slug": "six-open-source-ai-workflow-kits-you-can-actually-inspect", "title": "Six open-source AI workflow kits you can actually inspect", "summary": "Software Sausage released six open-source AI workflow kits, each containing a README, an editable evidence ledger, and a dependency-free shell verifier. All 17 verifiers pass at release v0.17.0, and the kits are designed to ensure that decision-critical evidence remains after AI tools complete their tasks. The workflows cover specification freezing, prompt regression testing, model evaluation, document conversion, browser trace retention, and dependency updates.", "body_md": "Disclosure: Software Sausage is our product and publishes these recipes. The kits are free and MIT-licensed. AI tools helped draft and edit this article; the workflow status and evidence limits are stated below.\n\nMost AI workflow lists answer the easiest question: which tools can be placed next to each other in a diagram?\n\nWe wanted to answer a harder one: what evidence should remain after the tools finish?\n\nThat produced six small workflow kits. Each one contains a README, an editable evidence ledger, and a dependency-free shell verifier. All 17 verifiers in the repository pass at release v0.17.0.\n\nThat is a structural claim, not a performance claim. The checks prove that the required files and fields exist. The new workflows remain explicitly marked “not benchmarked” until measured runs are published.\n\nUse GitHub Spec Kit to freeze the outcome, exclusions, acceptance criteria, and rollback boundary. Let one coding agent implement the reviewed tasks. Then run existing checks plus a small user-flow proof and ask a different model to compare the result with the original specification.\n\nThe artifact is not a generated plan. It is the linked specification, reviewed diff, executable checks, browser evidence, and rollback note.\n\nFreeze representative success, edge, and refusal cases before changing a prompt, model, tool, or instruction file. Promptfoo can run the baseline and candidate against the same cases while retaining assertions, latency, token use, cost, and failures.\n\nPrefer deterministic assertions before model grading. Also isolate the run: Promptfoo configurations can execute code and are not a sandbox.\n\nPick one real repository task and define hidden acceptance checks. Pin the harness, model endpoint, instructions, permissions, tools, and context budget. Run Qwen Code, Goose, OpenCode, or another candidate from fresh copies at least three times, then blind the labels before reviewing the artifacts.\n\nThe result should select a configuration for that job and environment, not declare a universal winner.\n\nRun MarkItDown and Docling on the same authorized documents. Score the raw output for ordering, tables, citations, omitted text, OCR errors, and usable source locations before giving it to a model.\n\nA fluent summary cannot repair a missing table cell. Reopen every decision-changing number, date, obligation, and citation in the source.\n\nFreeze the browser, viewport, data state, network conditions, and user action. Use Playwright to reproduce the flow and Chrome DevTools MCP to retain a trace, console output, and relevant network evidence. Make the smallest root-cause fix, then rerun the same conditions several times.\n\nOne lab trace is not field performance. Keep authenticated browser profiles away from an MCP client unless that access is deliberately required.\n\nLet Renovate propose a narrow update. Record the direct and transitive changes, lockfile diff, release notes, supported runtime range, and rollback version. Use OSV-Scanner and Semgrep as review inputs, then run the project's actual checks and a representative runtime flow.\n\nClean scans do not prove compatibility or the absence of vulnerabilities.\n\nEvery kit is available in the pinned [v0.17.0 release](https://github.com/Software-Sausage/recipes/releases/tag/v0.17.0). Run one on a disposable fixture. If its verifier passes while decision-critical evidence is missing, open an issue with the smallest safe reproduction. That is more useful than a star.\n\nThe readable library and the boundary for each workflow are in [Software Sausage](https://softwaresausage.com/blog/six-open-source-ai-workflow-kits?source=community&utm_campaign=dev_six_kits).\n\nPrimary references: [GitHub Spec Kit](https://github.com/github/spec-kit/blob/main/docs/index.md), [Promptfoo assertions](https://github.com/promptfoo/promptfoo/blob/main/site/docs/configuration/expected-outputs/index.md), [Qwen Code](https://github.com/QwenLM/qwen-code), [Goose](https://github.com/aaif-goose/goose), [Chrome DevTools MCP](https://github.com/ChromeDevTools/chrome-devtools-mcp), [MarkItDown MCP](https://github.com/microsoft/markitdown/blob/main/packages/markitdown-mcp/README.md), [Renovate](https://docs.renovatebot.com/), and [Semgrep CLI](https://semgrep.dev/docs/category/local-and-cli-scans).", "url": "https://wpnews.pro/news/six-open-source-ai-workflow-kits-you-can-actually-inspect", "canonical_source": "https://dev.to/softwaresausage/six-open-source-ai-workflow-kits-you-can-actually-inspect-amc", "published_at": "2026-09-03 22:20:44+00:00", "updated_at": "2026-09-03 22:55:01.978325+00:00", "lang": "en", "topics": ["ai-tools", "developer-tools", "ai-agents", "mlops"], "entities": ["Software Sausage", "GitHub Spec Kit", "Promptfoo", "Qwen Code", "Goose", "OpenCode", "MarkItDown", "Docling"], "alternates": {"html": "https://wpnews.pro/news/six-open-source-ai-workflow-kits-you-can-actually-inspect", "markdown": "https://wpnews.pro/news/six-open-source-ai-workflow-kits-you-can-actually-inspect.md", "text": "https://wpnews.pro/news/six-open-source-ai-workflow-kits-you-can-actually-inspect.txt", "jsonld": "https://wpnews.pro/news/six-open-source-ai-workflow-kits-you-can-actually-inspect.jsonld"}}