{"slug": "guardrails-in-the-prompt-aren-t-guardrails-an-authority-gate-for-claude-code", "title": "Guardrails in the Prompt Aren't Guardrails: An Authority Gate for Claude Code", "summary": "A developer built GODMODE V3, an authority and policy layer for Claude Code that moves agent guardrails out of the system prompt and into a PreToolUse hook that can deny any mutating tool call. The gate requires a live V3 authorization bound to a card, policy payload, repository, scope and exact 40-character source commit SHA, so changing a single byte of the payload causes the call to be denied. A GitHub Actions workflow runs the authority suite against the PR merge snapshot, reporting 11 passed with source_commit_sha=ad54bfd53e8880bce09a750870cc408d936e6ca4.", "body_md": "Most \"AI agent guardrails\" are text in a system prompt. The agent reads \"don't push to main\", and usually it complies.\n\n\"Usually\" isn't a control. I wanted the decision to sit **outside the model**, at the point where a tool call is about to change something, and to depend on state the model can't fake.\n\nThat's what **GODMODE V3** in `osa-execution-force-skills` does for Claude Code. It's a policy and authority layer that every mutating action has to pass through first.\n\nRepo: [github.com/HazEOskA/osa-execution-force-skills](https://github.com/HazEOskA/osa-execution-force-skills)\n\n```\nTASK\n  │\n  ▼\nGODMODE V3      resolve card + policy\n  │             bind repo / scope / payload / source SHA\n  ▼\nRuntimeV2       execute the downstream mission\n  │\n  ▼\nHOST ACTION\n  │\n  ▼\nEVIDENCE / VERIFICATION\n```\n\nThe system has two layers, and they are not peers:\n\nThe registry holds **150 operational cards across 14 domains**: backend, data, frontend, cloud/infra, AI/ML, security, SRE, Web3, low-level systems and more. Each card is structured data, not prose:\n\n`gdy`` twarde``nie`` minimalnyDowod``ryzyko`\nHere's a condensed version of card `H04` (cryptography in use). The repo is in Polish; I've translated it:\n\n```\n{\n  id: \"H04\",\n  name: \"Cryptography in use\",\n  when: \"passwords, encryption, signatures, tokens\",\n  hard: [\n    \"never implement your own primitives\",\n    \"passwords via argon2id or bcrypt, never SHA\",\n    \"AES-GCM or ChaCha20-Poly1305 with a unique nonce\",\n    \"constant-time comparison for secrets\",\n  ],\n  never: [\"encryption without authentication\", \"nonce from a counter reset on restart\"],\n  minimumProof: \"review of crypto library usage + verification of the randomness source\",\n  risk: \"R2\",\n}\n```\n\nThe `minimumProof` field matters most. A card doesn't just say how to do the work. It says what has to exist before anyone may claim it's done.\n\nClaude Code can run a hook before every tool call, and that hook can return `deny`. GODMODE wires three of them:\n\n```\n.claude/settings.json\n  ├─ PreToolUse   → V3 authority gate\n  ├─ PostToolUse  → authority/evidence state update\n  └─ Stop         → incomplete-flow stop guard\n```\n\nThe PreToolUse gate sorts every call into one of three kinds:\n\nA mutation must match a **live V3 authorization**. Here's how the RuntimeV2 path is checked (condensed):\n\n```\nif bare in RUNTIME_V2_REQUIRES_BOUND_PAYLOAD:\n    if v3_state.live_authorization(authority_state) is None:\n        return (\"RuntimeV2 execution is downstream-only. \"\n                \"Call osagm_authorize and obtain a live V3 authorization first.\")\n    if not v3_state.runtime_payload_matches(authority_state, tool_input):\n        return (\"osa_run_mission payload does not match the exact GODMODE \"\n                \"policy payload authorized for this session.\")\n```\n\nThe model can't just decide it's allowed. The authorization comes from `osagm_authorize`. It contains the card, the policy, the 40-character source commit SHA, the repository, the allowed scope, a task digest, and the **exact downstream payload plus its digest**. Change one byte of the payload and the call is denied.\n\nThe P0 gate exists to stop one specific failure: GODMODE code sits in the repo, but the host can still mutate through a different control plane. The enforced invariants:\n\n`source_commit_sha` must be an exact 40-character Git SHA. Missing or `UNKNOWN` provenance `§0`) returns `STOP`, not authority.\nThese are checked mechanically. A dedicated GitHub Actions workflow checks out the PR merge snapshot, starts the real V3 launcher and runs the authority suite:\n\n```\nV3_LAUNCHER_PROVENANCE_PASS\nsource_commit_sha=ad54bfd53e8880bce09a750870cc408d936e6ca4\ncard=A01\npolicy_payload=BOUND\n11 passed\n```\n\nI re-ran `tests/claude_hooks/test_v3_authority_wiring.py` on a fresh clone and got the same result: **11 passed**.\n\n```\nCLAIMED\n   ↓\nARTIFACT_PRESENT\n   ↓\nMECHANICALLY_VERIFIED\n   ↓\nINDEPENDENTLY_VERIFIED\n```\n\nA model saying \"tests passed\" is evidence of nothing. A mutation being *possible* doesn't prove the right authority path approved it. Every claim sits on one of these four rungs, and only the bottom two count as proof.\n\nFrom the README, deliberately:\n\nAgents can already write code. The open question is who signed off on *this* change, against *which* rule, on *which* commit, and what proves it was done. GODMODE is my answer for one host: a control path the model can't take over, and proof requirements on every card.\n\nAgents can execute. GODMODE controls the path and asks for proof.\n\nWhere does the authority decision live in your agent setup: in the prompt, or outside the model?", "url": "https://wpnews.pro/news/guardrails-in-the-prompt-aren-t-guardrails-an-authority-gate-for-claude-code", "canonical_source": "https://dev.to/hazeoska/guardrails-in-the-prompt-arent-guardrails-an-authority-gate-for-claude-code-52b5", "published_at": "2026-10-02 13:01:49+00:00", "updated_at": "2026-10-02 13:07:57.552170+00:00", "lang": "en", "topics": ["ai-agents", "ai-safety", "developer-tools", "ai-tools"], "entities": ["Claude Code", "GODMODE V3", "osa-execution-force-skills", "GitHub Actions", "osagm_authorize", "RuntimeV2"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/guardrails-in-the-prompt-aren-t-guardrails-an-authority-gate-for-claude-code", "markdown": "https://wpnews.pro/news/guardrails-in-the-prompt-aren-t-guardrails-an-authority-gate-for-claude-code.md", "text": "https://wpnews.pro/news/guardrails-in-the-prompt-aren-t-guardrails-an-authority-gate-for-claude-code.txt", "jsonld": "https://wpnews.pro/news/guardrails-in-the-prompt-aren-t-guardrails-an-authority-gate-for-claude-code.jsonld"}}