{"slug": "i-attack-tested-my-agent-s-seatbelt-here-s-what-survived", "title": "I attack-tested my agent's seatbelt. Here's what survived.", "summary": "A developer built a PreToolUse hook to block destructive Bash commands from an AI coding agent, but the initial blocklist failed four self-authored bypasses — base64 pipes, python3 -c, find -delete, and git clean -fdx — after passing its own five-case test suite. The second version replaces pattern matching with shlex-based parsing and a positive policy that allows destructive verbs only when their targets resolve inside the project root.", "body_md": "A blocklist for destructive agent commands passed its own tests — then failed four bypasses I wrote myself. The fix isn't smarter parsing. Tested receipts inside.\n\nI wrote a seatbelt for my AI coding agent: a `PreToolUse` hook that inspects every `Bash` call before it runs and blocks the destructive ones. It passed its own test suite — five runs, five green.\n\nThen I spent an evening attacking it. This post is the session log: what the blocklist caught, what walked straight through it, what the second version fixed, and where the wall actually is. Everything below ran for real; nothing is reconstructed.\n\nThe contract is simple: a `PreToolUse` hook receives the tool call as JSON on stdin, exits `0` to allow, `2` to block — and the hook's `stderr` goes back to the model, which reads why it was blocked and adapts. A skill is advice the model can weigh. A hook is an exit code.\n\nThe obvious implementation is a blocklist:\n\n``` bash\n#!/bin/sh\n# PreToolUse hook: reads { tool_name, tool_input } JSON on stdin.\n# exit 0 = allow | exit 2 = block, stderr goes back to the agent.\ncmd=$(python3 -c 'import json,sys; print(json.load(sys.stdin)[\"tool_input\"].get(\"command\",\"\"))' 2>/dev/null)\n\ncase \"$cmd\" in\n  *\"rm -rf \\$\"*|*\"rm -rf ~\"*|*\"rm -rf /\"*|\\\n  *\"git push --force\"*|*\"git push -f\"*|\\\n  *\"DROP TABLE\"*|*\"chmod -R 777 /\"*)\n    echo \"SEATBELT: blocked — destructive pattern matched in: $cmd\" >&2\n    exit 2 ;;\n  *)\n    exit 0 ;;\nesac\n```\n\nGreen across the board: `rm -rf $HOME/...` blocked, `mkdir -p build && ls` allowed, `git push --force origin main` blocked, `rm -rf /var/data/scratch` blocked, `DROP TABLE users` blocked. Five for five. Ship it?\n\nNote the `\\$` in the first pattern, by the way — the first draft didn't have it, because `sh` expands `$HOME` inside a `case` pattern. The guard was searching for an expanded path while the agent's command contained the literal text. The seatbelt failed its own test before I ever attacked it.\n\nThen I stopped testing it like a developer and started testing it like the thing it guards against:\n\n``` bash\n$ echo '{\"tool_input\":{\"command\":\"echo cm0gLXJmIC9EYXRhL3NlY3JldA== | base64 -d | sh\"}}' | guard-bash.sh\nexit=0\n\n$ echo '{\"tool_input\":{\"command\":\"python3 -c \\\"import shutil; shutil.rmtree(...)\\\"\"}}' | guard-bash.sh\nexit=0\n\n$ echo '{\"tool_input\":{\"command\":\"find / -name \\\"*.bak\\\" -delete\"}}' | guard-bash.sh\nexit=0\n\n$ echo '{\"tool_input\":{\"command\":\"git clean -fdx\"}}' | guard-bash.sh\nexit=0\n```\n\nFour attacks, four passes. A base64 pipe hides the verb entirely; `python3 -c` never spells `rm`; `find -delete` is deletion without the letter sequence; `git clean -fdx` is destruction wearing a porcelain face. The blocklist lost before it started, for a structural reason: the set of destructive commands is unbounded, and the shell is a language *designed* for composition. You cannot enumerate what you should fear.\n\nSo stop listing the bad commands and start stating the good territory. Positive policy: destructive verbs are allowed, but their targets must stay inside the project. Parsing gets real — `shlex`, not string matching:\n\n``` bash\n#!/usr/bin/env python3\n\"\"\"PreToolUse hook, v2 — destructive verbs must target paths inside PROJECT.\"\"\"\nimport json, shlex, sys, os\n\nPROJECT = os.path.realpath(os.environ.get(\"PROJECT_ROOT\", os.getcwd()))\n\ndef inside_project(path: str) -> bool:\n    p = os.path.realpath(path)\n    return p == PROJECT or p.startswith(PROJECT + os.sep)\n\ndef targets_of(argv: list[str]) -> list[str]:\n    verb = os.path.basename(argv[0])\n    if verb == \"rm\":\n        return [a for a in argv[1:] if not a.startswith(\"-\")]\n    if verb == \"git\" and len(argv) > 1 and argv[1] == \"clean\":\n        return argv[2:]\n    return []\n\ncall = json.load(sys.stdin)\ntry:\n    argv = shlex.split(call.get(\"tool_input\", {}).get(\"command\", \"\"))\nexcept ValueError:\n    print(\"SEATBELT: unparseable command — refusing\", file=sys.stderr)\n    sys.exit(2)\n\nbad = [t for t in targets_of(argv)\n       if t.startswith(\"/\") and not inside_project(t)]\nif bad:\n    print(f\"SEATBELT: destructive targets outside {PROJECT}: {bad}\", file=sys.stderr)\n    sys.exit(2)\nsys.exit(0)\n```\n\nAnd it failed its own test again — the good kind of failure. `rm -rf /tmp/seatbelt/build` was blocked as outside the project, because on macOS `/tmp` is a symlink to `/private/tmp`: `realpath` fixed the target while the project root stayed unprefixed. The fix is one line — `realpath` **both** sides — and it's the kind of bug you only find by running the guard against your own machine's pathologies.\n\nAfter the fix: `rm -rf /Data/secret` blocked, `git clean -fdx /etc` blocked, `rm -rf /tmp/seatbelt/build` allowed. Three for three.\n\nAnd still: `python3 -c \"import shutil; shutil.rmtree(...)\"` walks through, because it never spells a destructive verb the parser knows. v2 catches better than v1; it does not catch everything. Nothing that parses the command catches everything.\n\nThe layer that can't be parsed around is the one that doesn't read the command at all:\n\n``` bash\n$ chmod 555 wall/            # the parent directory loses write permission\n$ rm -rf wall/secret\nrm: wall/secret: Permission denied\nexit=1\n```\n\nA filesystem permission blocked the deletion with no parser, no pattern list, and nothing for a model to argue with. This is the unglamorous answer most agent-security threads circle past: the sandbox *is* the seatbelt — a separate OS user, a container, a scoped filesystem — and the hook is the polite, inspectable layer in front of it. [Railway's post-incident philosophy](https://blog.railway.com/p/your-ai-wants-to-nuke-your-database) says the same thing in product language: make the destructive thing slow, make the recoverable thing fast. And their [eval harness against destructive behavior](https://dev.to/reidmarlow/action-scaling-at-the-harness-boundary-beats-trajectory-re-runs-n5d) is the maintenance loop — a guard you don't re-attack on a schedule is a guard that's already rotting.\n\nThe final architecture is three layers, each with a different job: the **hook** as confirmation gate and tripwire — deterministic, inspectable, honest about patterns; the **OS** as the wall — permissions and sandboxing that no command string can charm; and the **attack session** as the maintenance loop — a calendar entry, not a vibe.\n\nTwo honest limits. Every layer described here except the filesystem one is a tripwire, and the python-one-liner class walks through all of them — which is exactly why the OS layer isn't optional. And this was one evening of attacks by the guard's own author; an adversary with more evenings will find more. The seatbelt isn't done. It's just honest about what it is: the layer that catches the boring, irreversible mistakes before the wall has to.\n\nWhat did your agent run last night that you couldn't have parsed?", "url": "https://wpnews.pro/news/i-attack-tested-my-agent-s-seatbelt-here-s-what-survived", "canonical_source": "https://dev.to/slabb/i-attack-tested-my-agents-seatbelt-heres-what-survived-15mj", "published_at": "2026-10-03 15:33:25+00:00", "updated_at": "2026-10-03 15:38:08.756845+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "ai-safety"], "entities": ["Bash", "Python", "shlex", "git"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-attack-tested-my-agent-s-seatbelt-here-s-what-survived", "markdown": "https://wpnews.pro/news/i-attack-tested-my-agent-s-seatbelt-here-s-what-survived.md", "text": "https://wpnews.pro/news/i-attack-tested-my-agent-s-seatbelt-here-s-what-survived.txt", "jsonld": "https://wpnews.pro/news/i-attack-tested-my-agent-s-seatbelt-here-s-what-survived.jsonld"}}