{"slug": "part-1-make-your-ai-agents-boring-the-determinism-layer", "title": "Part 1: Make your AI agents boring: the determinism layer", "summary": "A six-level maturity model for running LLM systems in production begins with a \"determinism layer\" that contains the model to a single node so the surrounding system is ordinary, testable code, according to a technical post on the approach. The design treats an agent as a pure function that returns a proposed decision with no authority to change business state, while a separate, heavily-tested \"substrate\" component applies proposals only after required approval. The post argues the constraint makes agents testable, safe, and composable, and replaces free-form ReAct loops with a fixed sequence of nodes where the model is used only where judgment is needed.", "body_md": "The trick to putting LLM agents in high-stakes systems isn’t a smarter model — it’s containing the model to one node so the rest of the system is ordinary, testable code. Here are the structural moves, with the contracts *and types to implement them.*\n\nDemos love autonomous agents that loop, call tools, and “figure it out.” Production hates them. The moment an agent’s behavior depends on which path the model wandered down today, you can’t test it, can’t audit it, and can’t let it touch anything that matters. In a regulated or high-consequence system — money movement, healthcare, infrastructure — “it usually works” is a non-starter.\n\nThis is Level 1 of a six-level maturity model for running LLM systems in production: **the determinism layer**. Before you can do evals, confidence routing, or anything else higher up the stack, you need the part underneath to hold still. The whole game at this level is one idea — **contain the model to a single node so everything around it is ordinary, testable code** — expressed through a handful of design moves. You keep the model’s intelligence and remove almost all of the unpredictability. This post gives the contracts, not just the concepts.\n\n### \n\nThe most important rule: **an agent is a pure function** from context to a *proposed* decision — pure with respect to business state, and deterministic given its gateway (the one nondeterministic call, the model, is injected so tests can fake it). It has no authority to change the world. A separate, dumb, heavily-tested component — the **substrate** — applies proposals, and only after the required approval.\n\n``` python\nfrom typing import Protocol, Literal\nfrom dataclasses import dataclass\n\n@dataclass(frozen=True)\nclass Proposal:\n    decision_id: str\n    capability: str\n    action: dict                  # the structured, proposed change — NOT yet applied\n    confidence: float\n    routing: Literal[\"auto\", \"hitl_recommended\", \"hitl_required\", \"reject\"]\n    reasoning: list[str]\n    evidence: list[dict]\n\nclass Agent(Protocol):\n    def propose(self, ctx: \"Context\") -> Proposal: ...     # pure: no side effects on business state\n\nclass Substrate(Protocol):\n    def apply(self, proposal: Proposal, approval: \"Approval\") -> \"Effect\": ...  # the ONLY mutator\n```\n\nBecause `propose() is pure, its headline property is one assertion:`\n\n``` python\ndef test_propose_is_pure():\n    agent = ClassifyAgent(gateway=FakeGateway(scripted))\n    assert agent.propose(ctx) == agent.propose(ctx)   # same context → same proposal; no world touched\n```\n\nThe boundary buys you three properties:\n\n- **Testable.**`propose()` is a pure function — same context, same proposal. No mocking the world to test the logic.\n- **Safe.** A jailbroken or buggy agent produces a bad*proposal* , not a bad*action* . The blast radius stops at “a guardrail or a human said no.”\n- **Composable.** Agents never call other agents. Work flows through the substrate (e.g. a cases table + a scheduler), so there are no hidden chains of side effects to reason about.\n\nThis single constraint turns “an AI did something we can’t explain” into “an AI *suggested* something, and here’s exactly what approved it.”\n\n### \n\nFree-form ReAct loops are great for exploration and terrible for guarantees. Model each capability as a **fixed sequence of nodes** where the model is used only where judgment is needed and everything else is ordinary code. The arrow sketch below is abridged for readability; the full node list (with `pre_check and` memory_write`) is in the next section.\n\n```\nentry → load context → reason (LLM) → output guardrail → verify →\n        judge (sampled) → compose confidence → route → prepare proposal → record → exit\nFree-form ReAct loop     Fixed-graph (this)\n--------------  -----------------------  ---------------------------------\nControl flow    model decides next step  known in advance\nTestable        hard (path varies)       each node in isolation\nLatency / cost  unbounded                bounded, predictable\nAudit           reconstruct from trace   uniform row every time\nRight for       open-ended exploration   decisions that must be guaranteed\n```\n\nYou give up some cleverness; you get back the ability to reason about what the system will do. The clever part — judgment on messy inputs — stays exactly where the model is good at it, in one node, surrounded by code you can read.\n\n### \n\n“Make agents deterministic” is hard to act on until you’ve seen the shape of one. So let’s walk a single request through the graph, node by node — the implementation behind the diagram above.\n\n**The state object.** Every node reads and writes one typed state value threaded through the graph. Making this explicit is half the battle — it’s what lets you test a node in isolation by constructing a state and asserting on the result.\n\n``` python\nfrom typing import TypedDict, Literal, Optional\n\nclass GraphState(TypedDict):\n    decision_id: str\n    tenant_id: str\n    identity: dict                      # who/what authority (set at entry)\n    inputs: dict                        # validated request\n    context: dict                       # loaded memory/reference slices\n    model_output: Optional[dict]        # raw structured output from the LLM node\n    guardrail: dict                     # blocks/redactions applied\n    verification: dict                  # deterministic check results\n    judge: Optional[dict]               # second-opinion result (if sampled)\n    confidence: Optional[float]         # composed score\n    routing: Optional[Literal[\"auto\", \"hitl_recommended\", \"hitl_required\", \"reject\"]]\n    proposal: Optional[dict]            # final shaped proposal\n    ledger_entry_id: Optional[str]\n```\n\n**The node contract.** Every node is the same shape: `state -> state`. Deterministic except the one model node. This uniformity is why you can unit-test each node and reason about the whole.\n\n``` python\nfrom typing import Protocol\n\nclass Node(Protocol):\n    name: str\n    def run(self, state: GraphState, deps: \"Deps\") -> GraphState: ...\n```\n\n**The graph.** Eight nodes are shared across every capability; only a few are capability-specific. New capability = implement ~4 nodes, inherit the other 8.\n\n```\n1  entry              mint decision_id, bind tenant + identity              [shared]\n2  pre_check          validate inputs, resolve refs, cheap early-outs       [capability]\n3  context_load       fetch only the context this decision needs            [shared]\n4  llm_decision       the reasoning step — structured in, structured out    [capability]\n5  output_guardrail   PII / policy scrub on the model output                [shared]\n6  verification       deterministic, capability-specific correctness checks [capability]\n7  judge              sampled second-model review (high-stakes)             [shared]\n8  confidence_compose composed score from the signals                       [shared]\n9  routing            auto vs hitl vs reject                                [shared]\n10 prepare_proposal   shape the final proposal                              [capability]\n11 memory_write       write the agent's own audit memory (not business state)[shared]\n12 exit               append the immutable ledger entry                     [shared]\n```\n\nThe two load-bearing nodes are `llm_decision and` verification`. The first is the only nondeterministic node — structured input, schema-constrained output, retry-on-mismatch — and everything around it treats its output as untrusted until checked (more on that next section). The second is deterministic, capability-specific code that the whole “contain the model” thesis rests on:\n\n``` python\ndef run(self, state, deps):\n    out = state[\"model_output\"]\n    checks = {\n        \"in_enum\":   out[\"decision\"] in ALLOWED_DECISIONS,        # can't return an off-list action\n        \"schema_ok\": matches_schema(out, DECISION_SCHEMA),\n        \"rules_ok\":  deps.rules.check(out, state[\"inputs\"]),      # business invariants\n    }\n    state[\"verification\"] = {\"checks\": checks, \"score\": sum(checks.values()) / len(checks)}\n    return state\n```\n\nThe point of the fixed shape: you can read the control flow (the path *is* the graph), the model stays contained to one node, and every decision is uniform — same shape every time means the same audit row every time, and the same place to add a check, a metric, or a guardrail.\n\n### \n\nNode 4 deserves its own treatment, because a surprising amount of LLM fragility comes from one choice: letting the model return free text and then parsing it. Prose is ambiguous, the format drifts between calls, and your downstream code is one unexpected phrasing away from breaking. “Sure! It’s probably Approve, though it could be Escalate” — now you own an NLU problem to extract `Approve`, and tomorrow's rephrase breaks your regex. You've coupled your system to the model's *prose style*, the least stable thing about it.\n\nThe fix is boring and powerful: **constrain the model to emit a validated structure**, and treat anything else as a failed call to retry. There are three ways to constrain, picked by how hard the guarantee must be:\n\n```\nmethod                                         how                             guarantee                             use when\n---------------------------------------------  ------------------------------  ------------------------------------  --------------------------------------\nSchema-guided (JSON Schema / response_format)  ask for JSON matching a schema  strong, provider-enforced             most cases\nTool / function call                           model emits a typed tool call   strong; natural for \"do X with args\"  the decision maps to an action\nGrammar-constrained decoding                   constrain tokens to a grammar   hard guarantee (can't emit invalid)   strict/regulated formats, local models\n# the contract as types — downstream consumes a typed object, never prose\nfrom pydantic import BaseModel\nfrom typing import Literal\n\nclass Decision(BaseModel):\n    decision: Literal[\"approve\", \"escalate\", \"reject\"]   # an enum is itself a guardrail\n    confidence: float\n    reasons: list[str]\n```\n\nConstrained generation reduces malformed output; it doesn’t eliminate it. So close the loop: validate every response, and on a mismatch retry — feeding the validation error back. Bound the retries and fail closed.\n\n``` python\nfrom pydantic import ValidationError\nclass NonConformingOutput(Exception): ...\n\ndef decide(base_prompt, model, max_retries=2) -> Decision:\n    prompt = base_prompt\n    for _ in range(max_retries + 1):\n        raw = model.generate(prompt, schema=Decision.model_json_schema())\n        try:\n            return Decision.model_validate_json(raw)            # success: typed object\n        except ValidationError as e:\n            # rebuild from base_prompt — don't append onto the already-appended prompt (it compounds)\n            prompt = f\"{base_prompt}\\n\\nYour previous output was invalid: {e}. Return JSON only.\"\n    raise NonConformingOutput()      # fail closed — never hand downstream a guess\n```\n\nValidation at the boundary means the rest of the system only ever sees well-formed decisions; the messy “did the model behave” question is contained to this one function. You trade a vague NLU problem for a crisp validation problem — no parsing layer, stability across model/prompt swaps, a built-in guardrail (a fixed enum *cannot* return something off-list), and a typed contract golden tests can assert against.\n\nOne caveat worth stating loudly: structure constrains *form*, not *correctness*. A perfectly valid `{\"decision\":\"approve\",\"confidence\":0.99}` can be completely wrong. Structured output removes the parsing failure mode, not the judgment failure mode — which is why it's the *floor* of a production system, under evals, verification, and confidence, not a substitute for them.\n\n### \n\nConfidence (model signal + verification + sampled judge) is composed at the confidence node; the routing node turns it into one of four paths with a single threshold, and `prepare_proposal` later stamps that result onto the proposal:\n\n``` php\ndef route(confidence: float, verified: bool, T: float = 0.85) -> str:  # T is per-capability config, not a constant\n    if not verified:           return \"reject\"\n    if confidence >= T:        return \"auto\"\n    if confidence >= T - 0.2:  return \"hitl_recommended\"   # close: pre-fill the proposal for a human\n    return \"hitl_required\"                                 # low: a human decides from scratch\n```\n\n`verified` is derived from the verification node (`verified = all(checks.values())`), not a field the agent sets on itself, and the router emits all four `routing` states. Note the order: routing runs *before* the proposal is shaped (node 9 → node 10), so it takes a plain confidence and a verified flag — not a `Proposal`. Start `T` conservative — everything to a human — and lower it per slice only as data proves it safe. Human attention, the expensive resource, gets spent exactly where the system is unsure.\n\n### \n\nEvery proposed decision becomes **one immutable row** — written as the last node of every decision, unconditionally. Not a log line; the canonical record of what happened and why.\n\n```\nCREATE TABLE decision_ledger (\n  decision_id     TEXT PRIMARY KEY,\n  ts              TIMESTAMPTZ NOT NULL,\n  tenant_id       TEXT NOT NULL,\n  capability      TEXT NOT NULL,\n  inputs_hash     TEXT NOT NULL,          -- hash, not the raw sensitive payload\n  model_version   TEXT NOT NULL,\n  prompt_version  TEXT NOT NULL,\n  decision        JSONB NOT NULL,\n  confidence      REAL NOT NULL,\n  routing         TEXT NOT NULL,          -- auto | hitl_* | reject\n  outcome         TEXT,                    -- recorded later as a NEW superseding row, never an in-place UPDATE\n  supersedes      TEXT REFERENCES decision_ledger(decision_id),\n  prev_hash       TEXT,                    -- optional hash-chain for tamper-evidence\n  entry_hash      TEXT\n);\n-- append-only: no UPDATE/DELETE grants; corrections are new rows that set `supersedes`.\n```\n\nHash the inputs, don’t warehouse them — verifiability without the liability. Make it append-only: corrections supersede, never overwrite. If you can `UPDATE` the ledger, it’s not an audit trail; revoke the grant. And never skip the write under load — it’s the record of record, not droppable telemetry.\n\n### \n\nIf you’ve followed all of the above — fixed graphs, the model contained to one node — you’ll eventually hit a problem that doesn’t fit: something open-ended where the model must look something up, reason about what it found, maybe look up more, then decide. That *is* what ReAct-style tool loops are for. The mistake isn’t the loop; it’s the *unbounded* loop. Allow autonomy as a deliberate, rail-guarded exception — and nowhere else.\n\nFour rails keep it controlled: **(1) a hard iteration cap** — `for, never` while(not done)`, so the model doesn’t decide when to stop; **(2) a per-capability tool allow-list** — it can only call tools on an explicit list scoped to this capability; **(3) a full per-iteration trace** — every step recorded, so the decision is replayable; and **(4) the same exits as everything else** — the loop’s *output* still flows through guardrails, verification, confidence, and routing, and still *proposes*, never acts.\n\n```\nclass DisallowedTool(Exception): ...    # raised when the loop reaches for an off-list tool\n\nALLOWED_TOOLS = {                       # rail 2: per-capability allow-list, not \"all tools\"\n    \"enrich_request\": {\"search_kb\", \"fetch_record\", \"lookup_reference\"},\n}\n\n@dataclass\nclass Step:                              # rail 3: one trace row per iteration\n    i: int; action: str; args: dict; result_digest: str\n\ndef bounded_react(state, deps, capability, MAX_STEPS=6) -> Proposal:\n    allow = ALLOWED_TOOLS[capability]\n    trace: list[Step] = []\n    for i in range(MAX_STEPS):                                   # rail 1: hard cap\n        action = deps.model.next_action(state, tools=sorted(allow))\n        if action.is_final:                                     # check FIRST — a final answer carries no tool\n            return finalize(action.proposal, trace)             # rail 4: still a PROPOSAL\n        if action.tool not in allow:                            # defense in depth\n            raise DisallowedTool(action.tool)\n        result = deps.tools[action.tool](**action.args)\n        trace.append(Step(i, action.tool, action.args, digest(result)))\n        state = state.with_observation(result)                  # loop-local state type, not the fixed-graph GraphState\n    return escalate(\"hit step cap\", trace)                      # bounded: cap hit → returns a Proposal routed \"hitl_required\"\n```\n\nThe `trace` rides into the ledger entry, so a loop-based decision is exactly as reconstructable as a fixed-graph one. The two rails worth a test each — it always terminates, and it can't reach an off-list tool:\n\n``` python\ndef test_always_terminates():\n    deps = fake_deps(model=never_final)                 # a model that never returns is_final\n    p = bounded_react(state, deps, \"enrich_request\", MAX_STEPS=3)\n    assert p.routing == \"hitl_required\"                 # hit the cap → routed to a human, didn't hang\n\ndef test_disallowed_tool_refused():\n    deps = fake_deps(model=calls(\"danger_tool\"))        # a tool not on the allow-list\n    with pytest.raises(DisallowedTool):\n        bounded_react(state, deps, \"enrich_request\")\n```\n\nThe decision rule for whether you even need this: *does reaching the decision require steps whose number and order depend on what’s discovered along the way?* No → fixed graph (most capabilities). Yes → bounded loop, four rails, documented as an exception. If you reach for a loop “to be safe” or “for flexibility,” stop — that’s usually the fixed-graph case in a costume. Flexibility you don’t need is nondeterminism you’ll debug at 2am.\n\n### \n\n- **The agent writes business state “just this once.”** Now it’s not pure, not testable, and a bug is an incident instead of a bad proposal. Keep the substrate the*only* mutator.\n- **Implicit state** passed as ad-hoc tuples/dicts between steps — you lose the ability to test a node in isolation. Make the state a typed object.\n- **“Return JSON” in the prompt with no schema enforcement** — you’ll still get prose, fences, or trailing commentary. Use schema/tool/grammar enforcement, then validate and retry; constrained ≠ guaranteed.\n- **“Confidence” lifted from the model.** Miscalibrated; compose it from independent signals instead.\n- **Mutable or skipped audit.** If you can`UPDATE` the ledger it's not an audit trail; if you drop the write under load you have no record of record.\n- **Unbounded loop “for flexibility.”** Default to the fixed graph. When you do loop, use a`for cap and a per-capability allow-list — never` while not done` or “all tools available.”\n\n### \n\nLevel 1 is one discipline applied consistently: contain the model to one node, make the agent a pure function that only *proposes*, constrain that node to a validated schema, compose confidence from independent signals to route the close calls to a human, let a dumb substrate be the sole mutator after approval, and write every decision to an append-only ledger. When a capability genuinely needs autonomy, budget it — a step cap, a tool allow-list, a trace, the same output checks as everything else.\n\nThe model still does what it’s uniquely good at — judgment on messy inputs — but the system around it is deterministic, testable, and auditable. Boring, in the best way: predictable enough to test, contained enough to trust, and defensible enough to ship. That’s the floor everything else in this series is built on.\n\n*Series: Running LLM systems in production — Level 1 of 6: Determinism.*", "url": "https://wpnews.pro/news/part-1-make-your-ai-agents-boring-the-determinism-layer", "canonical_source": "https://stackoverflow.blog/2026/10/07/part-1-make-your-ai-agents-boring-the-determinism-layer/", "published_at": "2026-10-07 20:14:49+00:00", "updated_at": "2026-10-07 20:46:23.921424+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "artificial-intelligence", "ai-safety", "mlops"], "entities": ["ReAct"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/part-1-make-your-ai-agents-boring-the-determinism-layer", "markdown": "https://wpnews.pro/news/part-1-make-your-ai-agents-boring-the-determinism-layer.md", "text": "https://wpnews.pro/news/part-1-make-your-ai-agents-boring-the-determinism-layer.txt", "jsonld": "https://wpnews.pro/news/part-1-make-your-ai-agents-boring-the-determinism-layer.jsonld"}}