I attack-tested my agent's seatbelt. Here's what survived. A developer built a PreToolUse hook to block destructive Bash commands from an AI coding agent, but the initial blocklist failed four self-authored bypasses — base64 pipes, python3 -c, find -delete, and git clean -fdx — after passing its own five-case test suite. The second version replaces pattern matching with shlex-based parsing and a positive policy that allows destructive verbs only when their targets resolve inside the project root. A blocklist for destructive agent commands passed its own tests — then failed four bypasses I wrote myself. The fix isn't smarter parsing. Tested receipts inside. I wrote a seatbelt for my AI coding agent: a PreToolUse hook that inspects every Bash call before it runs and blocks the destructive ones. It passed its own test suite — five runs, five green. Then I spent an evening attacking it. This post is the session log: what the blocklist caught, what walked straight through it, what the second version fixed, and where the wall actually is. Everything below ran for real; nothing is reconstructed. The contract is simple: a PreToolUse hook receives the tool call as JSON on stdin, exits 0 to allow, 2 to block — and the hook's stderr goes back to the model, which reads why it was blocked and adapts. A skill is advice the model can weigh. A hook is an exit code. The obvious implementation is a blocklist: bash /bin/sh PreToolUse hook: reads { tool name, tool input } JSON on stdin. exit 0 = allow | exit 2 = block, stderr goes back to the agent. cmd=$ python3 -c 'import json,sys; print json.load sys.stdin "tool input" .get "command","" ' 2 /dev/null case "$cmd" in "rm -rf \$" | "rm -rf ~" | "rm -rf /" |\ "git push --force" | "git push -f" |\ "DROP TABLE" | "chmod -R 777 /" echo "SEATBELT: blocked — destructive pattern matched in: $cmd" &2 exit 2 ;; exit 0 ;; esac Green across the board: rm -rf $HOME/... blocked, mkdir -p build && ls allowed, git push --force origin main blocked, rm -rf /var/data/scratch blocked, DROP TABLE users blocked. Five for five. Ship it? Note the \$ in the first pattern, by the way — the first draft didn't have it, because sh expands $HOME inside a case pattern. The guard was searching for an expanded path while the agent's command contained the literal text. The seatbelt failed its own test before I ever attacked it. Then I stopped testing it like a developer and started testing it like the thing it guards against: bash $ echo '{"tool input":{"command":"echo cm0gLXJmIC9EYXRhL3NlY3JldA== | base64 -d | sh"}}' | guard-bash.sh exit=0 $ echo '{"tool input":{"command":"python3 -c \"import shutil; shutil.rmtree ... \""}}' | guard-bash.sh exit=0 $ echo '{"tool input":{"command":"find / -name \" .bak\" -delete"}}' | guard-bash.sh exit=0 $ echo '{"tool input":{"command":"git clean -fdx"}}' | guard-bash.sh exit=0 Four attacks, four passes. A base64 pipe hides the verb entirely; python3 -c never spells rm ; find -delete is deletion without the letter sequence; git clean -fdx is destruction wearing a porcelain face. The blocklist lost before it started, for a structural reason: the set of destructive commands is unbounded, and the shell is a language designed for composition. You cannot enumerate what you should fear. So stop listing the bad commands and start stating the good territory. Positive policy: destructive verbs are allowed, but their targets must stay inside the project. Parsing gets real — shlex , not string matching: bash /usr/bin/env python3 """PreToolUse hook, v2 — destructive verbs must target paths inside PROJECT.""" import json, shlex, sys, os PROJECT = os.path.realpath os.environ.get "PROJECT ROOT", os.getcwd def inside project path: str - bool: p = os.path.realpath path return p == PROJECT or p.startswith PROJECT + os.sep def targets of argv: list str - list str : verb = os.path.basename argv 0 if verb == "rm": return a for a in argv 1: if not a.startswith "-" if verb == "git" and len argv 1 and argv 1 == "clean": return argv 2: return call = json.load sys.stdin try: argv = shlex.split call.get "tool input", {} .get "command", "" except ValueError: print "SEATBELT: unparseable command — refusing", file=sys.stderr sys.exit 2 bad = t for t in targets of argv if t.startswith "/" and not inside project t if bad: print f"SEATBELT: destructive targets outside {PROJECT}: {bad}", file=sys.stderr sys.exit 2 sys.exit 0 And it failed its own test again — the good kind of failure. rm -rf /tmp/seatbelt/build was blocked as outside the project, because on macOS /tmp is a symlink to /private/tmp : realpath fixed the target while the project root stayed unprefixed. The fix is one line — realpath both sides — and it's the kind of bug you only find by running the guard against your own machine's pathologies. After the fix: rm -rf /Data/secret blocked, git clean -fdx /etc blocked, rm -rf /tmp/seatbelt/build allowed. Three for three. And still: python3 -c "import shutil; shutil.rmtree ... " walks through, because it never spells a destructive verb the parser knows. v2 catches better than v1; it does not catch everything. Nothing that parses the command catches everything. The layer that can't be parsed around is the one that doesn't read the command at all: bash $ chmod 555 wall/ the parent directory loses write permission $ rm -rf wall/secret rm: wall/secret: Permission denied exit=1 A filesystem permission blocked the deletion with no parser, no pattern list, and nothing for a model to argue with. This is the unglamorous answer most agent-security threads circle past: the sandbox is the seatbelt — a separate OS user, a container, a scoped filesystem — and the hook is the polite, inspectable layer in front of it. Railway's post-incident philosophy https://blog.railway.com/p/your-ai-wants-to-nuke-your-database says the same thing in product language: make the destructive thing slow, make the recoverable thing fast. And their eval harness against destructive behavior https://dev.to/reidmarlow/action-scaling-at-the-harness-boundary-beats-trajectory-re-runs-n5d is the maintenance loop — a guard you don't re-attack on a schedule is a guard that's already rotting. The final architecture is three layers, each with a different job: the hook as confirmation gate and tripwire — deterministic, inspectable, honest about patterns; the OS as the wall — permissions and sandboxing that no command string can charm; and the attack session as the maintenance loop — a calendar entry, not a vibe. Two honest limits. Every layer described here except the filesystem one is a tripwire, and the python-one-liner class walks through all of them — which is exactly why the OS layer isn't optional. And this was one evening of attacks by the guard's own author; an adversary with more evenings will find more. The seatbelt isn't done. It's just honest about what it is: the layer that catches the boring, irreversible mistakes before the wall has to. What did your agent run last night that you couldn't have parsed?