Part 1: Make your AI agents boring: the determinism layer A six-level maturity model for running LLM systems in production begins with a "determinism layer" that contains the model to a single node so the surrounding system is ordinary, testable code, according to a technical post on the approach. The design treats an agent as a pure function that returns a proposed decision with no authority to change business state, while a separate, heavily-tested "substrate" component applies proposals only after required approval. The post argues the constraint makes agents testable, safe, and composable, and replaces free-form ReAct loops with a fixed sequence of nodes where the model is used only where judgment is needed. The trick to putting LLM agents in high-stakes systems isn’t a smarter model — it’s containing the model to one node so the rest of the system is ordinary, testable code. Here are the structural moves, with the contracts and types to implement them. Demos love autonomous agents that loop, call tools, and “figure it out.” Production hates them. The moment an agent’s behavior depends on which path the model wandered down today, you can’t test it, can’t audit it, and can’t let it touch anything that matters. In a regulated or high-consequence system — money movement, healthcare, infrastructure — “it usually works” is a non-starter. This is Level 1 of a six-level maturity model for running LLM systems in production: the determinism layer . Before you can do evals, confidence routing, or anything else higher up the stack, you need the part underneath to hold still. The whole game at this level is one idea — contain the model to a single node so everything around it is ordinary, testable code — expressed through a handful of design moves. You keep the model’s intelligence and remove almost all of the unpredictability. This post gives the contracts, not just the concepts. The most important rule: an agent is a pure function from context to a proposed decision — pure with respect to business state, and deterministic given its gateway the one nondeterministic call, the model, is injected so tests can fake it . It has no authority to change the world. A separate, dumb, heavily-tested component — the substrate — applies proposals, and only after the required approval. python from typing import Protocol, Literal from dataclasses import dataclass @dataclass frozen=True class Proposal: decision id: str capability: str action: dict the structured, proposed change — NOT yet applied confidence: float routing: Literal "auto", "hitl recommended", "hitl required", "reject" reasoning: list str evidence: list dict class Agent Protocol : def propose self, ctx: "Context" - Proposal: ... pure: no side effects on business state class Substrate Protocol : def apply self, proposal: Proposal, approval: "Approval" - "Effect": ... the ONLY mutator Because propose is pure, its headline property is one assertion: python def test propose is pure : agent = ClassifyAgent gateway=FakeGateway scripted assert agent.propose ctx == agent.propose ctx same context → same proposal; no world touched The boundary buys you three properties: - Testable. propose is a pure function — same context, same proposal. No mocking the world to test the logic. - Safe. A jailbroken or buggy agent produces a bad proposal , not a bad action . The blast radius stops at “a guardrail or a human said no.” - Composable. Agents never call other agents. Work flows through the substrate e.g. a cases table + a scheduler , so there are no hidden chains of side effects to reason about. This single constraint turns “an AI did something we can’t explain” into “an AI suggested something, and here’s exactly what approved it.” Free-form ReAct loops are great for exploration and terrible for guarantees. Model each capability as a fixed sequence of nodes where the model is used only where judgment is needed and everything else is ordinary code. The arrow sketch below is abridged for readability; the full node list with pre check and memory write is in the next section. entry → load context → reason LLM → output guardrail → verify → judge sampled → compose confidence → route → prepare proposal → record → exit Free-form ReAct loop Fixed-graph this -------------- ----------------------- --------------------------------- Control flow model decides next step known in advance Testable hard path varies each node in isolation Latency / cost unbounded bounded, predictable Audit reconstruct from trace uniform row every time Right for open-ended exploration decisions that must be guaranteed You give up some cleverness; you get back the ability to reason about what the system will do. The clever part — judgment on messy inputs — stays exactly where the model is good at it, in one node, surrounded by code you can read. “Make agents deterministic” is hard to act on until you’ve seen the shape of one. So let’s walk a single request through the graph, node by node — the implementation behind the diagram above. The state object. Every node reads and writes one typed state value threaded through the graph. Making this explicit is half the battle — it’s what lets you test a node in isolation by constructing a state and asserting on the result. python from typing import TypedDict, Literal, Optional class GraphState TypedDict : decision id: str tenant id: str identity: dict who/what authority set at entry inputs: dict validated request context: dict loaded memory/reference slices model output: Optional dict raw structured output from the LLM node guardrail: dict blocks/redactions applied verification: dict deterministic check results judge: Optional dict second-opinion result if sampled confidence: Optional float composed score routing: Optional Literal "auto", "hitl recommended", "hitl required", "reject" proposal: Optional dict final shaped proposal ledger entry id: Optional str The node contract. Every node is the same shape: state - state . Deterministic except the one model node. This uniformity is why you can unit-test each node and reason about the whole. python from typing import Protocol class Node Protocol : name: str def run self, state: GraphState, deps: "Deps" - GraphState: ... The graph. Eight nodes are shared across every capability; only a few are capability-specific. New capability = implement ~4 nodes, inherit the other 8. 1 entry mint decision id, bind tenant + identity shared 2 pre check validate inputs, resolve refs, cheap early-outs capability 3 context load fetch only the context this decision needs shared 4 llm decision the reasoning step — structured in, structured out capability 5 output guardrail PII / policy scrub on the model output shared 6 verification deterministic, capability-specific correctness checks capability 7 judge sampled second-model review high-stakes shared 8 confidence compose composed score from the signals shared 9 routing auto vs hitl vs reject shared 10 prepare proposal shape the final proposal capability 11 memory write write the agent's own audit memory not business state shared 12 exit append the immutable ledger entry shared The two load-bearing nodes are llm decision and verification . The first is the only nondeterministic node — structured input, schema-constrained output, retry-on-mismatch — and everything around it treats its output as untrusted until checked more on that next section . The second is deterministic, capability-specific code that the whole “contain the model” thesis rests on: python def run self, state, deps : out = state "model output" checks = { "in enum": out "decision" in ALLOWED DECISIONS, can't return an off-list action "schema ok": matches schema out, DECISION SCHEMA , "rules ok": deps.rules.check out, state "inputs" , business invariants } state "verification" = {"checks": checks, "score": sum checks.values / len checks } return state The point of the fixed shape: you can read the control flow the path is the graph , the model stays contained to one node, and every decision is uniform — same shape every time means the same audit row every time, and the same place to add a check, a metric, or a guardrail. Node 4 deserves its own treatment, because a surprising amount of LLM fragility comes from one choice: letting the model return free text and then parsing it. Prose is ambiguous, the format drifts between calls, and your downstream code is one unexpected phrasing away from breaking. “Sure It’s probably Approve, though it could be Escalate” — now you own an NLU problem to extract Approve , and tomorrow's rephrase breaks your regex. You've coupled your system to the model's prose style , the least stable thing about it. The fix is boring and powerful: constrain the model to emit a validated structure , and treat anything else as a failed call to retry. There are three ways to constrain, picked by how hard the guarantee must be: method how guarantee use when --------------------------------------------- ------------------------------ ------------------------------------ -------------------------------------- Schema-guided JSON Schema / response format ask for JSON matching a schema strong, provider-enforced most cases Tool / function call model emits a typed tool call strong; natural for "do X with args" the decision maps to an action Grammar-constrained decoding constrain tokens to a grammar hard guarantee can't emit invalid strict/regulated formats, local models the contract as types — downstream consumes a typed object, never prose from pydantic import BaseModel from typing import Literal class Decision BaseModel : decision: Literal "approve", "escalate", "reject" an enum is itself a guardrail confidence: float reasons: list str Constrained generation reduces malformed output; it doesn’t eliminate it. So close the loop: validate every response, and on a mismatch retry — feeding the validation error back. Bound the retries and fail closed. python from pydantic import ValidationError class NonConformingOutput Exception : ... def decide base prompt, model, max retries=2 - Decision: prompt = base prompt for in range max retries + 1 : raw = model.generate prompt, schema=Decision.model json schema try: return Decision.model validate json raw success: typed object except ValidationError as e: rebuild from base prompt — don't append onto the already-appended prompt it compounds prompt = f"{base prompt}\n\nYour previous output was invalid: {e}. Return JSON only." raise NonConformingOutput fail closed — never hand downstream a guess Validation at the boundary means the rest of the system only ever sees well-formed decisions; the messy “did the model behave” question is contained to this one function. You trade a vague NLU problem for a crisp validation problem — no parsing layer, stability across model/prompt swaps, a built-in guardrail a fixed enum cannot return something off-list , and a typed contract golden tests can assert against. One caveat worth stating loudly: structure constrains form , not correctness . A perfectly valid {"decision":"approve","confidence":0.99} can be completely wrong. Structured output removes the parsing failure mode, not the judgment failure mode — which is why it's the floor of a production system, under evals, verification, and confidence, not a substitute for them. Confidence model signal + verification + sampled judge is composed at the confidence node; the routing node turns it into one of four paths with a single threshold, and prepare proposal later stamps that result onto the proposal: php def route confidence: float, verified: bool, T: float = 0.85 - str: T is per-capability config, not a constant if not verified: return "reject" if confidence = T: return "auto" if confidence = T - 0.2: return "hitl recommended" close: pre-fill the proposal for a human return "hitl required" low: a human decides from scratch verified is derived from the verification node verified = all checks.values , not a field the agent sets on itself, and the router emits all four routing states. Note the order: routing runs before the proposal is shaped node 9 → node 10 , so it takes a plain confidence and a verified flag — not a Proposal . Start T conservative — everything to a human — and lower it per slice only as data proves it safe. Human attention, the expensive resource, gets spent exactly where the system is unsure. Every proposed decision becomes one immutable row — written as the last node of every decision, unconditionally. Not a log line; the canonical record of what happened and why. CREATE TABLE decision ledger decision id TEXT PRIMARY KEY, ts TIMESTAMPTZ NOT NULL, tenant id TEXT NOT NULL, capability TEXT NOT NULL, inputs hash TEXT NOT NULL, -- hash, not the raw sensitive payload model version TEXT NOT NULL, prompt version TEXT NOT NULL, decision JSONB NOT NULL, confidence REAL NOT NULL, routing TEXT NOT NULL, -- auto | hitl | reject outcome TEXT, -- recorded later as a NEW superseding row, never an in-place UPDATE supersedes TEXT REFERENCES decision ledger decision id , prev hash TEXT, -- optional hash-chain for tamper-evidence entry hash TEXT ; -- append-only: no UPDATE/DELETE grants; corrections are new rows that set supersedes . Hash the inputs, don’t warehouse them — verifiability without the liability. Make it append-only: corrections supersede, never overwrite. If you can UPDATE the ledger, it’s not an audit trail; revoke the grant. And never skip the write under load — it’s the record of record, not droppable telemetry. If you’ve followed all of the above — fixed graphs, the model contained to one node — you’ll eventually hit a problem that doesn’t fit: something open-ended where the model must look something up, reason about what it found, maybe look up more, then decide. That is what ReAct-style tool loops are for. The mistake isn’t the loop; it’s the unbounded loop. Allow autonomy as a deliberate, rail-guarded exception — and nowhere else. Four rails keep it controlled: 1 a hard iteration cap — for, never while not done , so the model doesn’t decide when to stop; 2 a per-capability tool allow-list — it can only call tools on an explicit list scoped to this capability; 3 a full per-iteration trace — every step recorded, so the decision is replayable; and 4 the same exits as everything else — the loop’s output still flows through guardrails, verification, confidence, and routing, and still proposes , never acts. class DisallowedTool Exception : ... raised when the loop reaches for an off-list tool ALLOWED TOOLS = { rail 2: per-capability allow-list, not "all tools" "enrich request": {"search kb", "fetch record", "lookup reference"}, } @dataclass class Step: rail 3: one trace row per iteration i: int; action: str; args: dict; result digest: str def bounded react state, deps, capability, MAX STEPS=6 - Proposal: allow = ALLOWED TOOLS capability trace: list Step = for i in range MAX STEPS : rail 1: hard cap action = deps.model.next action state, tools=sorted allow if action.is final: check FIRST — a final answer carries no tool return finalize action.proposal, trace rail 4: still a PROPOSAL if action.tool not in allow: defense in depth raise DisallowedTool action.tool result = deps.tools action.tool action.args trace.append Step i, action.tool, action.args, digest result state = state.with observation result loop-local state type, not the fixed-graph GraphState return escalate "hit step cap", trace bounded: cap hit → returns a Proposal routed "hitl required" The trace rides into the ledger entry, so a loop-based decision is exactly as reconstructable as a fixed-graph one. The two rails worth a test each — it always terminates, and it can't reach an off-list tool: python def test always terminates : deps = fake deps model=never final a model that never returns is final p = bounded react state, deps, "enrich request", MAX STEPS=3 assert p.routing == "hitl required" hit the cap → routed to a human, didn't hang def test disallowed tool refused : deps = fake deps model=calls "danger tool" a tool not on the allow-list with pytest.raises DisallowedTool : bounded react state, deps, "enrich request" The decision rule for whether you even need this: does reaching the decision require steps whose number and order depend on what’s discovered along the way? No → fixed graph most capabilities . Yes → bounded loop, four rails, documented as an exception. If you reach for a loop “to be safe” or “for flexibility,” stop — that’s usually the fixed-graph case in a costume. Flexibility you don’t need is nondeterminism you’ll debug at 2am. - The agent writes business state “just this once.” Now it’s not pure, not testable, and a bug is an incident instead of a bad proposal. Keep the substrate the only mutator. - Implicit state passed as ad-hoc tuples/dicts between steps — you lose the ability to test a node in isolation. Make the state a typed object. - “Return JSON” in the prompt with no schema enforcement — you’ll still get prose, fences, or trailing commentary. Use schema/tool/grammar enforcement, then validate and retry; constrained ≠ guaranteed. - “Confidence” lifted from the model. Miscalibrated; compose it from independent signals instead. - Mutable or skipped audit. If you can UPDATE the ledger it's not an audit trail; if you drop the write under load you have no record of record. - Unbounded loop “for flexibility.” Default to the fixed graph. When you do loop, use a for cap and a per-capability allow-list — never while not done or “all tools available.” Level 1 is one discipline applied consistently: contain the model to one node, make the agent a pure function that only proposes , constrain that node to a validated schema, compose confidence from independent signals to route the close calls to a human, let a dumb substrate be the sole mutator after approval, and write every decision to an append-only ledger. When a capability genuinely needs autonomy, budget it — a step cap, a tool allow-list, a trace, the same output checks as everything else. The model still does what it’s uniquely good at — judgment on messy inputs — but the system around it is deterministic, testable, and auditable. Boring, in the best way: predictable enough to test, contained enough to trust, and defensible enough to ship. That’s the floor everything else in this series is built on. Series: Running LLM systems in production — Level 1 of 6: Determinism.