cd /news/ai-agents/guardrails-in-the-prompt-aren-t-guar… · home › topics › ai-agents › article
[ARTICLE · art-143881] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Guardrails in the Prompt Aren't Guardrails: An Authority Gate for Claude Code

A developer built GODMODE V3, an authority and policy layer for Claude Code that moves agent guardrails out of the system prompt and into a PreToolUse hook that can deny any mutating tool call. The gate requires a live V3 authorization bound to a card, policy payload, repository, scope and exact 40-character source commit SHA, so changing a single byte of the payload causes the call to be denied. A GitHub Actions workflow runs the authority suite against the PR merge snapshot, reporting 11 passed with source_commit_sha=ad54bfd53e8880bce09a750870cc408d936e6ca4.

by read3 min views3 publishedOct 2, 2026

Most "AI agent guardrails" are text in a system prompt. The agent reads "don't push to main", and usually it complies.

"Usually" isn't a control. I wanted the decision to sit outside the model, at the point where a tool call is about to change something, and to depend on state the model can't fake.

That's what GODMODE V3 in osa-execution-force-skills does for Claude Code. It's a policy and authority layer that every mutating action has to pass through first.

Repo: github.com/HazEOskA/osa-execution-force-skills

TASK
  │
  ▼
GODMODE V3      resolve card + policy
  │             bind repo / scope / payload / source SHA
  ▼
RuntimeV2       execute the downstream mission
  │
  ▼
HOST ACTION
  │
  ▼
EVIDENCE / VERIFICATION

The system has two layers, and they are not peers:

The registry holds 150 operational cards across 14 domains: backend, data, frontend, cloud/infra, AI/ML, security, SRE, Web3, low-level systems and more. Each card is structured data, not prose:

gdy`` twarde``nie`` minimalnyDowod``ryzyko Here's a condensed version of card H04 (cryptography in use). The repo is in Polish; I've translated it:

{
  id: "H04",
  name: "Cryptography in use",
  when: "passwords, encryption, signatures, tokens",
  hard: [
    "never implement your own primitives",
    "passwords via argon2id or bcrypt, never SHA",
    "AES-GCM or ChaCha20-Poly1305 with a unique nonce",
    "constant-time comparison for secrets",
  ],
  never: ["encryption without authentication", "nonce from a counter reset on restart"],
  minimumProof: "review of crypto library usage + verification of the randomness source",
  risk: "R2",
}

The minimumProof field matters most. A card doesn't just say how to do the work. It says what has to exist before anyone may claim it's done.

Claude Code can run a hook before every tool call, and that hook can return deny. GODMODE wires three of them:

.claude/settings.json
  ├─ PreToolUse   → V3 authority gate
  ├─ PostToolUse  → authority/evidence state update
  └─ Stop         → incomplete-flow stop guard

The PreToolUse gate sorts every call into one of three kinds:

A mutation must match a live V3 authorization. Here's how the RuntimeV2 path is checked (condensed):

if bare in RUNTIME_V2_REQUIRES_BOUND_PAYLOAD:
    if v3_state.live_authorization(authority_state) is None:
        return ("RuntimeV2 execution is downstream-only. "
                "Call osagm_authorize and obtain a live V3 authorization first.")
    if not v3_state.runtime_payload_matches(authority_state, tool_input):
        return ("osa_run_mission payload does not match the exact GODMODE "
                "policy payload authorized for this session.")

The model can't just decide it's allowed. The authorization comes from osagm_authorize. It contains the card, the policy, the 40-character source commit SHA, the repository, the allowed scope, a task digest, and the exact downstream payload plus its digest. Change one byte of the payload and the call is denied.

The P0 gate exists to stop one specific failure: GODMODE code sits in the repo, but the host can still mutate through a different control plane. The enforced invariants:

source_commit_sha must be an exact 40-character Git SHA. Missing or UNKNOWN provenance §0) returns STOP, not authority. These are checked mechanically. A dedicated GitHub Actions workflow checks out the PR merge snapshot, starts the real V3 launcher and runs the authority suite:

V3_LAUNCHER_PROVENANCE_PASS
source_commit_sha=ad54bfd53e8880bce09a750870cc408d936e6ca4
card=A01
policy_payload=BOUND
11 passed

I re-ran tests/claude_hooks/test_v3_authority_wiring.py on a fresh clone and got the same result: 11 passed.

CLAIMED
   ↓
ARTIFACT_PRESENT
   ↓
MECHANICALLY_VERIFIED
   ↓
INDEPENDENTLY_VERIFIED

A model saying "tests passed" is evidence of nothing. A mutation being possible doesn't prove the right authority path approved it. Every claim sits on one of these four rungs, and only the bottom two count as proof.

From the README, deliberately:

Agents can already write code. The open question is who signed off on this change, against which rule, on which commit, and what proves it was done. GODMODE is my answer for one host: a control path the model can't take over, and proof requirements on every card.

Agents can execute. GODMODE controls the path and asks for proof.

Where does the authority decision live in your agent setup: in the prompt, or outside the model?

── more in #ai-agents 4 stories · sorted by recency
── more on @claude code 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/guardrails-in-the-pr…] indexed:0 read:3min 2026-10-02 · —