cd /news/ai-agents/i-attack-tested-my-agent-s-seatbelt-… · home › topics › ai-agents › article
[ARTICLE · art-144500] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

I attack-tested my agent's seatbelt. Here's what survived.

A developer built a PreToolUse hook to block destructive Bash commands from an AI coding agent, but the initial blocklist failed four self-authored bypasses — base64 pipes, python3 -c, find -delete, and git clean -fdx — after passing its own five-case test suite. The second version replaces pattern matching with shlex-based parsing and a positive policy that allows destructive verbs only when their targets resolve inside the project root.

by read5 min views1 publishedOct 3, 2026

A blocklist for destructive agent commands passed its own tests — then failed four bypasses I wrote myself. The fix isn't smarter parsing. Tested receipts inside.

I wrote a seatbelt for my AI coding agent: a PreToolUse hook that inspects every Bash call before it runs and blocks the destructive ones. It passed its own test suite — five runs, five green.

Then I spent an evening attacking it. This post is the session log: what the blocklist caught, what walked straight through it, what the second version fixed, and where the wall actually is. Everything below ran for real; nothing is reconstructed.

The contract is simple: a PreToolUse hook receives the tool call as JSON on stdin, exits 0 to allow, 2 to block — and the hook's stderr goes back to the model, which reads why it was blocked and adapts. A skill is advice the model can weigh. A hook is an exit code.

The obvious implementation is a blocklist:

#!/bin/sh
cmd=$(python3 -c 'import json,sys; print(json.load(sys.stdin)["tool_input"].get("command",""))' 2>/dev/null)

case "$cmd" in
  *"rm -rf \$"*|*"rm -rf ~"*|*"rm -rf /"*|\
  *"git push --force"*|*"git push -f"*|\
  *"DROP TABLE"*|*"chmod -R 777 /"*)
    echo "SEATBELT: blocked — destructive pattern matched in: $cmd" >&2
    exit 2 ;;
  *)
    exit 0 ;;
esac

Green across the board: rm -rf $HOME/... blocked, mkdir -p build && ls allowed, git push --force origin main blocked, rm -rf /var/data/scratch blocked, DROP TABLE users blocked. Five for five. Ship it?

Note the \$ in the first pattern, by the way — the first draft didn't have it, because sh expands $HOME inside a case pattern. The guard was searching for an expanded path while the agent's command contained the literal text. The seatbelt failed its own test before I ever attacked it.

Then I stopped testing it like a developer and started testing it like the thing it guards against:

$ echo '{"tool_input":{"command":"echo cm0gLXJmIC9EYXRhL3NlY3JldA== | base64 -d | sh"}}' | guard-bash.sh
exit=0

$ echo '{"tool_input":{"command":"python3 -c \"import shutil; shutil.rmtree(...)\""}}' | guard-bash.sh
exit=0

$ echo '{"tool_input":{"command":"find / -name \"*.bak\" -delete"}}' | guard-bash.sh
exit=0

$ echo '{"tool_input":{"command":"git clean -fdx"}}' | guard-bash.sh
exit=0

Four attacks, four passes. A base64 pipe hides the verb entirely; python3 -c never spells rm; find -delete is deletion without the letter sequence; git clean -fdx is destruction wearing a porcelain face. The blocklist lost before it started, for a structural reason: the set of destructive commands is unbounded, and the shell is a language designed for composition. You cannot enumerate what you should fear.

So stop listing the bad commands and start stating the good territory. Positive policy: destructive verbs are allowed, but their targets must stay inside the project. Parsing gets real — shlex, not string matching:

#!/usr/bin/env python3
"""PreToolUse hook, v2 — destructive verbs must target paths inside PROJECT."""
import json, shlex, sys, os

PROJECT = os.path.realpath(os.environ.get("PROJECT_ROOT", os.getcwd()))

def inside_project(path: str) -> bool:
    p = os.path.realpath(path)
    return p == PROJECT or p.startswith(PROJECT + os.sep)

def targets_of(argv: list[str]) -> list[str]:
    verb = os.path.basename(argv[0])
    if verb == "rm":
        return [a for a in argv[1:] if not a.startswith("-")]
    if verb == "git" and len(argv) > 1 and argv[1] == "clean":
        return argv[2:]
    return []

call = json.load(sys.stdin)
try:
    argv = shlex.split(call.get("tool_input", {}).get("command", ""))
except ValueError:
    print("SEATBELT: unparseable command — refusing", file=sys.stderr)
    sys.exit(2)

bad = [t for t in targets_of(argv)
       if t.startswith("/") and not inside_project(t)]
if bad:
    print(f"SEATBELT: destructive targets outside {PROJECT}: {bad}", file=sys.stderr)
    sys.exit(2)
sys.exit(0)

And it failed its own test again — the good kind of failure. rm -rf /tmp/seatbelt/build was blocked as outside the project, because on macOS /tmp is a symlink to /private/tmp: realpath fixed the target while the project root stayed unprefixed. The fix is one line — realpath both sides — and it's the kind of bug you only find by running the guard against your own machine's pathologies.

After the fix: rm -rf /Data/secret blocked, git clean -fdx /etc blocked, rm -rf /tmp/seatbelt/build allowed. Three for three.

And still: python3 -c "import shutil; shutil.rmtree(...)" walks through, because it never spells a destructive verb the parser knows. v2 catches better than v1; it does not catch everything. Nothing that parses the command catches everything.

The layer that can't be parsed around is the one that doesn't read the command at all:

$ chmod 555 wall/            # the parent directory loses write permission
$ rm -rf wall/secret
rm: wall/secret: Permission denied
exit=1

A filesystem permission blocked the deletion with no parser, no pattern list, and nothing for a model to argue with. This is the unglamorous answer most agent-security threads circle past: the sandbox is the seatbelt — a separate OS user, a container, a scoped filesystem — and the hook is the polite, inspectable layer in front of it. Railway's post-incident philosophy says the same thing in product language: make the destructive thing slow, make the recoverable thing fast. And their eval harness against destructive behavior is the maintenance loop — a guard you don't re-attack on a schedule is a guard that's already rotting.

The final architecture is three layers, each with a different job: the hook as confirmation gate and tripwire — deterministic, inspectable, honest about patterns; the OS as the wall — permissions and sandboxing that no command string can charm; and the attack session as the maintenance loop — a calendar entry, not a vibe.

Two honest limits. Every layer described here except the filesystem one is a tripwire, and the python-one-liner class walks through all of them — which is exactly why the OS layer isn't optional. And this was one evening of attacks by the guard's own author; an adversary with more evenings will find more. The seatbelt isn't done. It's just honest about what it is: the layer that catches the boring, irreversible mistakes before the wall has to.

What did your agent run last night that you couldn't have parsed?

── more in #ai-agents 4 stories · sorted by recency
── more on @bash 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/i-attack-tested-my-a…] indexed:0 read:5min 2026-10-03 · —