Ducklab – dev harness that built itself: 416 runs, $176, local models Ducklab, an Apache-2.0 self-hosted software development harness built with a Go engine, CLI, and desktop app, has completed 416 runs at a total cost of $176 using local models first, and has developed itself through 111 accepted tasks. The harness, designed for multi-LLM use, ensures verdicts are based on exit codes rather than model opinions, and supports local models via llama.cpp and vLLM alongside OpenAI-compatible and Anthropic endpoints. It operates over MCP and records all decisions, with a gate that reproduces tests from clean checkouts before any commit is accepted. A full-cycle software development harness that is multi-LLM by default and honest by construction. In one block: self-hosted development harness Go engine + CLI + desktop, Linux first · brief → requirements → spec → plan → build → review → release · verdicts are exit codes, never model opinions · local models first llama.cpp, vLLM beside any OpenAI-compatible or Anthropic endpoint · operable by humans or by other agents over MCP with recorded, attributed decisions · Apache-2.0 · develops itself the run records in .ducklab/ are the receipts . Agents: start at AGENTS.md /jrullan/ducklab/blob/main/AGENTS.md and . /jrullan/ducklab/blob/main/llms.txt llms.txt You give it a brief. It writes requirements, a spec and a plan; builds tasks with one model or several arguing; runs your project's real test gate; and stops for you before anything is committed. Every model call is logged. No model ever decides a verdict. /jrullan/ducklab/blob/main/docs/screenshots/council.gif A real council intake, recorded live and sped up: the architect streams the draft, a different model reviews it, the budget ticks in cents — and the run stops at your gate. Total cost of what you just watched: $0.07. It was built for local models first — the two that built most of it are a vLLM box on the LAN and a llama.cpp server on localhost, both priced at zero — and hosted models sit beside them in the same roster, measured by the same evidence. Most agentic coding tools assume one strong model and trust it. Ducklab assumes several cheap models and trusts none of them : The gate decides, never a model. A verdict is a command's exit code. A test-first run measures a green baseline before any test is written, the red over the new test after, and every accept reproduces the gate from a clean checkout of the committed sha — nothing lands that did not reproduce, and an accept whose reproduction fails takes its own commit back. Decorrelation everywhere. A different model reviews; a reviewer never learns who wrote the code absent from the payload, not hidden in the UI ; tournament judges choose blind; council critics read the draft, not each other. Work is a contract. A task's deliverables are the implementer's numbered checklist; it reports on each by number, the reviewer checks each against the diff, and an undelivered item summons the rubber duck — an advisor seat that wakes only on measured distress brake refusals, failure streaks, red gates and answers none , a note that sends the implementer straight back to work, or stop . Seats are chosen on evidence. Every duckling carries a scorecard — in-seat pass rate from your own runs, cost per run, coding index — and the roster board suggests seats from it, with the ranking criteria yours to reorder. Suggestions are rare and justified: pass rates rank by their Wilson lower bound, three runs minimum, locals never win on a $0 price. Nothing is unbounded. Turns, tokens, cost, wallclock, tool output, shell commands — every ceiling visible and liftable mid-run, on the record. Your documentation is not bounded by the model's window. Attach a wiki to a stage and a big seat reads it whole; a small seat gets each document digested to fit, the full text one ref read call away, and the gate names any document nobody opened. A 32k local model can be briefed by a quarter-million characters of reference material — the harness carries the working memory. /jrullan/ducklab/blob/main/docs/screenshots/runs.png The record does not round up: every run with its verdict, its cost, and whether its accept reproduced green from a clean checkout. And the existence proof: ducklab is developed inside ducklab. The plan, the bugs, the releases and 111 accepted tasks went through its own loop, driven by the same local and hosted models it measures — most recent features the escalation suggestions, the acceptance receipts, the MCPB release packaging, the multimodal chat were built by the duck, gated by a person. Don't take the claim on faith: git clone https://github.com/jrullan/ducklab && cd ducklab go build -o ducklab-cli ./cmd/ducklab for r in .ducklab/runs/ /receipt.json; do ./ducklab-cli proof verify "$r"; done Receipts ship with every accept since v0.7.0: the committed sha, the gate command, its exit code, and the clean-checkout reproduction verdict — facts a third party re-derives, never assessments. v0.7.0 , moving fast — seven releases in the first three weeks. Seven stages, five modes, the roster board with evidence and suggestions, reference documents with automatic digestion, skills managed from the desktop, a seated consultant you can chat with images included, vision verified before they are sent , bugs with screenshot evidence, adopt surveys with a deterministic coverage check at the gate, provider-aware queueing that says why a run waits, escalation suggestions when a seat measurably hits its ceiling, exportable acceptance receipts with ducklab proof verify , releases, autopilot, a CLI, a desktop app, and an MCP server — in the official MCP registry https://registry.modelcontextprotocol.io as io.github.jrullan/ducklab — that lets another model operate the whole loop with recorded, attributed decisions. docs/status.md /jrullan/ducklab/blob/main/docs/status.md tracks all acceptance criteria and does not round up. Where code and spec differ, the difference is recorded in . /jrullan/ducklab/blob/main/docs/decisions docs/decisions/ Needs Go 1.25+, Node 22+ for the desktop, and git. The CLI and engine are pure Go. The desktop is a Wails v3 app and needs the GTK/WebKit development packages: sudo apt install libgtk-3-dev libwebkit2gtk-4.1-dev Debian/Ubuntu names make desktop && make install On Ubuntu 24.04+ the desktop also needs an AppArmor profile — see decision 0003 /jrullan/ducklab/blob/main/docs/decisions/0003-apparmor-userns.md and packaging/apparmor/ . xcode-select --install the desktop build links against WebKit brew install go node make desktop && make install Honesty note: ducklab is developed and exercised daily on Linux. The CLI and engine compile-check for darwin/arm64 on every make cross , but no desktop build has been verified on a Mac yet — the first person to try it is the test, and make install gives you the CLI and engine either way. Please report whatever breaks. make install installs to ~/.local/bin — make sure it is on your PATH . It warns when the desktop binary predates frontend/src , because it will happily install a stale one. To exercise the frontend against the lightweight fake engine, run the engine and Vite in separate terminals, then open the browser with its connection details: go run ./cmd/fake-engine --port 8787 --token fake-token npm run dev --prefix frontend open http://localhost:5173/?engine=http://127.0.0.1:8787&token=fake-token The engine and token query parameters are available only in Vite dev builds. They can also be supplied as VITE DUCKLAB ENGINE and VITE DUCKLAB TOKEN environment variables. The desktop shell continues to use its injected window.ducklab connection. | What it is | | |---|---| ducklab-engine | The daemon. Owns every run. Binds 127.0.0.1 only, bearer token rotated each start. | ducklab | The CLI client. Holds no state; it asks the engine. | ducklab-desktop | The desktop app. Also a client, also holds no state. Starts or adopts the engine itself. | Provider keys come from the engine's environment at call time — export them before it starts, or launch the desktop through a wrapper that loads them from your keyring. The app tells you when the engine it adopted is missing a key this app has, with the restart button beside the words. From the desktop: Projects → New project , then Cycle → Draft it . From a terminal: cd ~/dev/myproject git init ducklab needs a git repo ducklab project init --name MyProject auto-starts the engine if none is running ducklab intake --from brief.txt brief → requirements ducklab spec requirements → spec ducklab plan spec → milestones and tasks ducklab run T-001 build it ducklab run accept r-20260729-... commit it ducklab review T-001 read the commit ducklab release plan --bump minor what shipped Each stage writes a .proposed file first and waits for you. accept promotes it; reject restores exactly what the run wrote and nothing else; "request changes" sends any draft — spec, plan, release notes — back with your note. Nothing is committed without you or without the autonomy level you explicitly granted . Reference documents ride any stage: --ref ~/wiki/product/ or the attach door in the desktop loads files or whole directories as background for the architect — grounded by two rules the prompt states outright: the approved requirements own the scope, and where a reference and the code disagree, the code is the truth. When the corpus outgrows the seat's context, each document is digested once cached by content hash , the full text stays reachable through the ref read tool, and the proposal card lists any document no seat ever opened. Adopting an existing codebase works the same way: intake reads the code and writes as-built requirements, the spec marks its sections as-built , and the plan stays deliberately empty — new work then enters through bug reports and plan amendments, which is how ducklab itself is developed. Your project declares its own truth in .ducklab/project.toml : the gate verify — with link deps and setup for what a clean checkout needs , how the app launches run with a preflight , and how the project's own binaries are rebuilt install so the whole loop runs without leaving ducklab. Gate and shell process trees always receive DUCKLAB RUN ID and DUCKLAB PROJECT ID . For example, excercise-tracker can use DATABASE URL=test db ${DUCKLAB RUN ID} in verify .tests , and a compose preflight can use ${DUCKLAB PROJECT ID} as its per-run project name. Ducklab guarantees identity only; provisioning and teardown remain the project's. ducklab provider set openrouter --url https://openrouter.ai/api/v1 \ --key-env OPENROUTER API KEY ducklab duckling set pato-sonnet --provider openrouter \ --model anthropic/claude-sonnet-4.5 \ --roles reviewer,judge --context 200000 \ --cost-in 3.0 --cost-out 15.0 ducklab duckling test pato-sonnet --prompt "say OK" --key-env is the name of an environment variable, never a key. No key is written to config, sent over the API, or kept in shell history. /jrullan/ducklab/blob/main/docs/screenshots/roster.png Seats are argued with evidence: pass rates from your own runs, cost per run, coding index — suggestions justified, never imposed. The desktop's Roster view is where seats are assigned: drag from the Flock onto a mode's seat, globally or per project, with each duckling's evidence on the card and the engine's suggestions beside the seats. Coding / intelligence / agentic indices come from OpenRouter's benchmarks endpoint when a duckling lives there; your own runs supply the rest. /jrullan/ducklab/blob/main/docs/screenshots/council-run.png The same machinery on real work: a council revising ducklab's own spec, 4.5M tokens in, paused once on a budget it asked to lift. ducklab run T-001 --mode