Build a Fail-Closed Ship Card Before an AI CLI Leaves Dry-Run A developer published a fail-closed "ship card" template for AI CLI tools that refuses to run unless committed evidence — dry-run stdout, a diff patch, and a zero exit code — is present alongside path locks, blast-radius caps, and a rollback plan. The accompanying Python checker validates a JSON card's required fields, rejects empty gates, non-integer or negative change limits, and missing or empty evidence files, exiting with code 2 on any failure. The author frames the script as a starter template rather than a sandbox product, noting it does not call a network or score a model. I stare at a green dry-run and still hesitate to ship. One passing fixture is not a write-mode signal for me. Would you let that helper touch a shared branch tonight? A solo CLI can look finished and still be unsafe to run. The missing piece is saved evidence, not another clever prompt. I keep a ship card beside the tool for that pause. The card is a small JSON file committed next to the CLI. It names gates, proof files, and a stop rule. If any proof is missing, the process exits. I am not asking the model to bless itself. I am asking a local script to refuse the run. Can a missing proof file count as a feature here? This is a template you can copy tonight. It is not a benchmark from my laptop. Fill every field from a run you actually saved. Name every directory the CLI is allowed to read. Name every file the CLI is allowed to write. Anything outside those lists is out of scope. I keep the write list painfully short on purpose. A helper that may touch the whole repo is not ready. Why hand it the whole tree this early? Run the helper once with writes turned off. Save stdout, the patch, and the exit code. Put those three files under an evidence directory. An empty evidence directory means you stop immediately. A chat log is not a patch you can review. Did you actually save the diff this time? Set a maximum changed-file count on the card. Set a maximum added-line count beside it. Refuse deletes unless the card explicitly allows them. A surprise rename should fail this gate immediately. A secret-shaped string in the patch should fail it too. Would you merge that change while half asleep? Write the exact command that restores the tree. Write the condition that abandons the whole experiment. Write what you do if the runner is gone. No exit plan means the card is still incomplete. I would rather shelve the branch than improvise. Shipping without rollback is how tiny tools linger. | Gate | Evidence you commit | Fail closed when | |---|---|---| | Path lock | allowed reads and writes | a list is empty, or a patch path is outside | | Dry-run | stdout, diff.patch, exit code | any file is missing or the code is not zero | | Blast radius | max files, max lines, delete flag | caps are blank, or a delete appears | | Exit plan | rollback command, abandon rule | either string is empty or untested | Read the row before you argue with the script. The script is the boring coworker in this workflow. Are you willing to lose an argument to it? Skip a step and the card becomes theater. A skipped step turns the later review into mush. Which step do you usually skip under time pressure? This script is a starter, not a sandbox product. It checks presence, a zero exit, deletes, and write paths. It does not call a network, and it does not score a model. bash /usr/bin/env python3 """Fail-closed ship card checker. Template, not a benchmark.""" import json import sys from pathlib import Path REQUIRED = "allowed reads", "allowed writes", "evidence dir", "max changed files", "max added lines", "allow deletes", "rollback cmd", "abandon if", def fail msg: str - None: print f"SHIP CARD FAIL: {msg}" raise SystemExit 2 def main - None: card path = Path sys.argv 1 if len sys.argv 1 else "ship-card.json" if not card path.is file : fail f"missing card: {card path}" card = json.loads card path.read text encoding="utf-8" for key in REQUIRED: if key not in card or card key in "", , None : fail f"empty gate: {key}" for key in "max changed files", "max added lines" : if not isinstance card key , int or card key < 0: fail f"{key} must be a non-negative integer" evidence = Path card "evidence dir" needed = "stdout.txt", "diff.patch", "exit code.txt" for name in needed: file = evidence / name if not file.is file or file.stat .st size == 0: fail f"missing evidence: {file}" code = evidence / "exit code.txt" .read text encoding="utf-8" .strip if code = "0": fail f"dry-run exit was {code}" diff = evidence / "diff.patch" .read text encoding="utf-8" if "diff --git" not in diff: fail "diff.patch has no git diff header" if card "allow deletes" is not False and card "allow deletes" is not True: fail "allow deletes must be a boolean" if not card "allow deletes" and "\ndeleted file mode " in f"\n{diff}": fail "delete found but allow deletes is false" writes = card "allowed writes" for line in diff.splitlines : if not line.startswith "+++ b/" : continue path = line 6: allowed = path in writes or any path.startswith prefix.rstrip "/" + "/" for prefix in writes if not allowed: fail f"write outside allowlist: {path}" print "SHIP CARD OK" if name == " main ": main { "allowed reads": "src/", "tests/fixtures/" , "allowed writes": "src/format.py" , "evidence dir": "evidence/dry-run-01", "max changed files": 1, "max added lines": 40, "allow deletes": false, "rollback cmd": "git checkout -- src/format.py", "abandon if": "runner missing, empty diff, or secret-like token in patch" } The numeric caps must be integers, or the checker stops. It still does not count lines inside the patch. Add that counter before you trust write mode. Make the evidence folder, then run the checker. A missing patch should end with exit code 2. Do not catch that error and keep going. mkdir -p evidence/dry-run-01 python3 ship card.py ship-card.json echo $? Capture a real dry-run with the commands below. Adjust the CLI name to your own tool. Keep writes off until the card is green. python3 rewrite cli.py --dry-run src/format.py evidence/dry-run-01/stdout.txt printf '%s\n' "$?" evidence/dry-run-01/exit code.txt git diff -- src/format.py evidence/dry-run-01/diff.patch Save the CLI status before any later command runs. Otherwise you will record git, not the helper. rewrite cli.py is only a placeholder name. Use a bad patch when you want a guaranteed red run. This one deletes a file the card should protect. Your card should reject it if deletes are forbidden. diff --git a/src/format.py b/src/format.py deleted file mode 100644 index 1111111..0000000 --- a/src/format.py +++ /dev/null @@ -1,3 +0,0 @@ -def normalize text : - return text.strip Point evidence dir at that folder and rerun the checker. Confirm the printed failure and the non-zero exit. Then point the card back at the good folder. Also search the patch before you widen writes. The starter script does not scan for secrets. A green card can still hide a leaked key. rg -n "AKIA|sk-|BEGIN PRIVATE" evidence || echo "no obvious token" Add that search to your own local wrapper. I would not call the card done without it. Did the bad fixture fail loudly on your machine? Disclosure: This article was prepared as part of MonkeyCode's product outreach. I treat MonkeyCode free model access as one optional dry-run runner. A free server option can host that small canary off your laptop. I will not invent a token cap, a model name, or a hardware tier. Those two availability claims are only useful after you read them today. A stale screenshot is a failed evidence gate. Quotas, names, and duration can change without this article updating. If you use that path, add only fields you personally verified. Leave them blank when the page will not load. Blank means abandon the write, not guess a limit. The checker never calls MonkeyCode on your behalf. You still run your own command on the machine. { "model access": "free-model-access", "compute": "free-server-option", "limits checked on": "YYYY-MM-DD", "limits source": "account page reviewed today" } The card only records that the limits were looked up. Remove the product name and the four gates still stand. Path locks do not need a vendor to matter. If MonkeyCode is the runner you picked, open the account page first. Paste only the free-model and free-server facts you can see today. Then let the checker argue, not your memory. Give this experiment one evening, not an open-ended week. I would stop after two red fixtures in a row. More retries usually mean the task itself is muddy. Practice the rollback before you actually need it. Dirty status after rollback is an abandon signal. Do not continue on a tree you cannot explain. git status --short git checkout -- src/format.py git status --short If the second status is not clean, shelve the branch. Write the reason into the abandon if field. Tomorrow you should not have to guess why it stopped. Do not use this card on customer data, payments, or auth code. Do not use it as a security audit or a compliance pack. Do not point it at secrets or production credentials. Skip it when your team already has a required change tool. A second checklist will rot within a week. Also skip it if you cannot store a local diff. Free model access can be wrong, slow, or simply unavailable. A free server option can vanish during the canary. This card does not promise uptime, capacity, or a lasting free tier. The path check is plain string matching, not a real sandbox. A tricky relative path can still slip through it. Treat the allowlist as a seatbelt, not a vault. Line caps and file caps are stored, not counted, in this version. That is a real gap, not a footnote to ignore. Add counters before write mode, or stay in dry-run. I am not publishing timings, prices, or model comparisons here. I do not have a fresh primary measurement to cite. If you need numbers, measure them yourself and date the note. Copy the checker, the sample card, and the bad fixture. Run the red case first, then one dry-run of your own CLI. Stop when a required field is still empty. What evidence file is still missing before you allow writes? That gap is the next build, not a new feature.