Part 4: Safety and governance for LLM systems: guardrails, PII, audit, and memory A four-part maturity model for LLM systems reaches Level 4, safety and governance, which requires layered guardrails that fail closed, PII handling at every boundary, an immutable audit trail, and memory scoped by category so one customer's data cannot reach another's decision. The model specifies six guardrail checkpoints — input/pre-prompt, grounding constraints, output scrub, verification, judge, and composed confidence with routing — each implemented as a small testable contract with a GuardrailResult action of "ok", "block", "redact", or "flag", and mandates that an errored or unavailable guardrail blocks or escalates rather than passing the request through. Every block is logged as a first-class signal via a guardrail_blocks_total metric labeled by layer and a bounded rule enum, never by matched content. The level where an LLM system stops being a demo and earns the right to touch real data and real decisions: layered guardrails that fail closed, PII handled at the boundary, an immutable audit trail, and scoped memory. By the time an LLM system is making decisions that matter, “it usually works” is no longer the bar. This is Level 4 of the maturity model — safety and governance — and it’s where four disciplines that teams tend to bolt on late have to be designed in instead. They share one idea: don’t trust a single point to do the right thing. Layer independent guardrails so a miss at one is caught at the next. Handle sensitive data at every boundary so it never accumulates where it shouldn’t. Write an immutable audit trail so “why did it do that?” has an answer months later. And scope memory by category so one customer’s data can never reach another’s decision — with a clean line between what you ship and what the system earns. None of these is exotic. Each is a small, testable contract enforced in code. Here is how they fit together. A lot of teams “add guardrails” by bolting one moderation filter onto the model output and calling it done. That’s the AI equivalent of a single firewall rule. Real safety, like real security, is defense in depth : several independent layers, each catching a different class of problem, arranged so a miss at one is caught at the next. Make every guardrail the same small contract, so you can add, remove, test, and reorder them independently. python from typing import Protocol, Literal from dataclasses import dataclass, field @dataclass class GuardrailResult: action: Literal "ok", "block", "redact", "flag" layer: str rule: str = "" detail: dict = field default factory=dict e.g. {"fields": "email" } class Guardrail Protocol : layer: str def check self, ctx: "Context" - GuardrailResult: ... A request then hits checkpoints on the way in and out: Layer Catches Stage - ----------------------------- ------------------------------------------------------ --------------------- 1 Input / pre-prompt injection, oversized/malformed input before the model 2 Grounding constraints model using disallowed actions/data; off-schema output shapes the model call 3 Output scrub policy violations, PII echoed back after generation 4 Verification business-rule / consistency violations deterministic code 5 Judge "plausible but wrong" that passed mechanical checks sampled second model 6 Composed confidence + routing the unknowns — anything still uncertain final net in ─▶ 1 input ─▶ 2 grounding ─▶ model ─▶ 3 output scrub ─▶ 4 verify ─▶ 5 judge ─▶ 6 confidence/route ─▶ act|escalate The order and the pass/fail logic live in one readable place. Two rules: fail closed an errored or unavailable guardrail blocks or escalates — it never waves the request through and log every block as a first-class signal. php def run guardrails ctx, layers, metrics - list GuardrailResult : applied = accumulate — a redacted result must be visible to the caller/audit for g in layers: try: r = g.check ctx except GuardrailError: NARROW — don't swallow your own bugs as a "block" metrics.incr "guardrail blocks total", layer=g.layer, rule="error" rule = bounded enum, NEVER matched content return applied + GuardrailResult "block", g.layer, "error" same label the metric emitted if r.action == "ok": continue metrics.incr "guardrail blocks total", layer=g.layer, rule=r.rule if r.action in "redact", "flag" : ctx.apply r ; applied.append r ; continue if r.action == "block": return applied + r stop at first hard block return applied + GuardrailResult "block", g.layer, "unknown action" unknown action == block fail closed return applied or GuardrailResult "ok", "all" The one test that matters most here proves it fails closed : python def test fail closed on error : class Boom: a guardrail that throws layer = "x" def check self, ctx : raise GuardrailError "boom" result = run guardrails ctx, Boom , NullMetrics assert result -1 .action == "block" an erroring guardrail BLOCKS, never silently passes Why layers beat one big filter: there’s no single point of failure injection slips past layer 1? grounding limits what it can do; an off output? scrub or verification catches it . Each layer is simple and testable — six single-purpose checks each have a clear contract, where one mega-filter is impossible to reason about. And different layers catch different failure modes: i nput scrub catches attacks, verification catches logic errors, the judge catches plausible-wrongness, confidence catches unknown unknowns. No single mechanism covers all four. One more reason to meter every block: guardrail blocks total{layer, rule} is one of the sharpest production health signals you have. A spike is an attack, a regression, or a bad deploy — page on it. The anti-patterns are mostly the inverse of the rules. Fail-open is worse than no guardrail — it gives false assurance. One layer doing everything is unmaintainable. Guardrails the model can talk past “please don’t do X” in the prompt aren’t guardrails — constrain the output space enum/schema so the disallowed thing is unrepresentable. Silent blocks mean you can’t tell an attack from a bug. And note that last layer-3 job — scrubbing PII the model echoed back — which is the natural handoff to the next discipline. AI systems are unusually hungry for data — they want rich context to reason well, and they generate records of everything they decide. That collides with a basic obligation: don’t accumulate sensitive personal data you don’t need. The resolution is to handle PII at the boundary — scrub it on the way in and out, and never let raw sensitive data settle into your stores, logs, or model traffic. Drive redaction from a declared classification, not ad-hoc if field == "email" scattered around. class Sensitivity Enum : PUBLIC = 0 ok anywhere INTERNAL = 1 ok in tier-1/2 logs + ledger summary PII = 2 mask in summaries; never in shared memory; tier-3 only if retained at all SECRET = 3 never persisted, never logged, never to the model unless essential FIELD POLICY = { the single source of truth "name": Sensitivity.PII, "email": Sensitivity.PII, "card number": Sensitivity.SECRET, "amount": Sensitivity.INTERNAL, "category": Sensitivity.PUBLIC, } Every place data moves between components is a boundary — into the system, into the model, into the ledger, into logs, out to a service. At each, ask: does what crosses here need raw sensitive data? Usually no — redact, mask, or hash before it crosses. inbound ─▶ scrub ─▶ working set ─▶ scrub ─▶ model ├─▶ redact + hash ─▶ ledger no raw PII └─▶ redact ─▶ logs no raw PII The ledger boundary deserves special care: you need to prove what a decision was made on, not retain the sensitive payload. Store a hash plus a redacted summary; re-hash later to prove equivalence — verifiability without the liability. php def classify key - Sensitivity: return FIELD POLICY.get key, Sensitivity.PII DEFAULT-DENY: unknown keys are treated as PII def redact obj : MUST recurse — real payloads are nested if isinstance obj, dict : return {k: "•••" if classify k .value = Sensitivity.PII.value else redact v for k, v in obj.items } if isinstance obj, list : return redact v for v in obj return obj def drop secrets obj : SECRET fields never even enter the hash input if isinstance obj, dict : return {k: drop secrets v for k, v in obj.items if classify k is not Sensitivity.SECRET} if isinstance obj, list : return drop secrets v for v in obj return obj def ledger view raw: dict, tenant key: bytes - dict: the ONLY way to write to the ledger safe = drop secrets raw HMAC with a per-tenant key, NOT bare sha256 — a plain hash of low-entropy PII card, email, phone is brute-forceable / rainbow-tableable, so "verifiability without liability" needs a key. return {"inputs hmac": "hmac-sha256:" + hmac sha256 tenant key, canonical json safe , "inputs summary": redact raw } Two decisions to pin once and reuse everywhere: a keyed hash HMAC, per-tenant key in KMS for any low-entropy field — a bare SHA-256 of a 16-digit number is reversible; b a single canonical-JSON spec e.g. RFC 8785 / JCS — sorted keys, normalized unicode, fixed number format , because the audit-ledger hash chain depends on re-hashing producing identical bytes. The model is an external boundary too — often a third party. Strip sensitive fields it doesn’t need to reason out of prompts, and scrub its output before persisting or returning, because models echo input back. That output scrub is exactly layer 3 above. The failure mode is “we redact in most places” — a leak with extra steps. Make scrubbing the only path through each boundary, so a developer can’t forget: ledger.append ledger view raw, tenant key there is no ledger.append raw — redaction isn't optional log.tier2 redact event the logging helper redacts; raw logging isn't exposed And design for residency and access up front. Sensitive data carries constraints on where it may live and who may read it — per-tenant keys, region-pinned storage for regulated tiers, access-gated audit reads. Retrofitting after data has spread everywhere is the nightmare you’re avoiding. The subtle leak isn’t out , it’s sideways — per-user context reaching the components that decide for other users — but that one is best handled structurally, in the memory model below. The first time someone asks “why did the system decide that for case X back in March?”, you find out whether you built an audit trail or just have logs. Logs are for debugging — they rotate, they’re unstructured, they’re not the truth. An audit ledger is the canonical, append-only, tamper-evident record of every decision and why. For anything consequential it’s not optional. CREATE TABLE decision ledger decision id TEXT PRIMARY KEY, -- threads through logs/traces for this decision ts TIMESTAMPTZ NOT NULL, tenant id TEXT NOT NULL, identity TEXT NOT NULL, -- who/what authority it ran with capability TEXT NOT NULL, inputs hash TEXT NOT NULL, -- keyed hash of canonical inputs NOT the raw payload inputs summary JSONB NOT NULL, -- redacted, PII-free human-readable summary model version TEXT NOT NULL, prompt version TEXT NOT NULL, decision JSONB NOT NULL, -- the structured decision confidence REAL, -- nullable: a manual supersede carries no model confidence routing TEXT NOT NULL, -- auto | hitl recommended | hitl required | reject | abstain outcome JSONB, -- the ONE mutable column, filled in later; EXCLUDED from the hash supersedes TEXT REFERENCES decision ledger decision id , seq BIGSERIAL, -- monotonic, for the hash chain prev hash TEXT, -- entry hash of seq-1 entry hash TEXT NOT NULL -- H canonical row-without-hash + prev hash ; -- App role gets INSERT + SELECT only. Revoke mutation in the DB, not just in code. REVOKE UPDATE, DELETE, TRUNCATE ON decision ledger FROM app role; GRANT UPDATE outcome ON decision ledger TO app role; -- column-level: ONLY outcome is writable later -- NOTE: this stops the APP, not a DB superuser/owner who can still DROP/TRUNCATE/rewrite. True -- append-only against an operator needs WORM/object-lock storage or external anchoring below . Three design choices do the heavy lifting. Append-only — corrections supersede, never overwrite. You never edit or delete an entry; a correction is a new row that points at the old one. seq 1041 decision=approve conf=0.91 supersedes=NULL seq 1207 decision=reject conf=NULL supersedes=