Part 3: Knowing when your agent doesn’t know: the confidence layer A six-level maturity model for running LLM systems in production places "confidence" at Level 3, arguing that an agent's most important output is a calibrated certainty score rather than its answer. The model composes that score from four independent signals — model self-report (weight 0.25), deterministic verification (0.35), an independent judge (0.25) and historical slice accuracy (0.15) — and drops the judge term with renormalization when a decision is unjudged, since only about 8% of decisions are judged. It instructs teams to validate the score with a reliability diagram, citing an example where a 0.90–1.00 predicted bucket was actually correct only 71% of the time, indicating the automation threshold is too loose. The most important number an agent produces isn’t its answer — it’s how sure it is. Compose that number from independent signals, check it’s calibrated, grade the high-stakes calls with a second model, and route the rest to humans well. This is Level 3 of a six-level maturity model for running LLM systems in production. Levels 1 and 2 got you to where the system works and you can see it working. Level 3 is confidence : the system acts on its own only when its calibrated confidence is high, grades the decisions that matter with an independent judge, and routes everything it’s unsure about to a human in a way that actually catches errors. These three things are one mechanism. Confidence is the dial that decides what runs automatically. A judge is one of the signals that dial is built from — and the trigger that pulls a decision off the automated path. The human handoff is where the dial sends everything below the line. Get all three right and you have a system that knows when it doesn’t know , and can therefore be trusted with more. Ask most LLM systems “how sure are you?” and you get nothing useful — either no number, or the model’s self-reported confidence, which is notoriously miscalibrated models are cheerfully certain when wrong . So you have to engineer a confidence signal. A production score is composed from independent signals, never lifted from the model’s own claim: @dataclass class Signals: model self: float the model's own score — useful but weighted DOWN overconfident verification: float fraction of deterministic checks passed 0..1 judge agreed: bool | None True/False if this decision was judged; None if not sampled most aren't historical: float this slice's past accuracy e.g. 0.99 vs 0.80 WEIGHTS = {"model self": 0.25, "verification": 0.35, "judge": 0.25, "historical": 0.15} def compose s: Signals - float: only ~8% of decisions are judged. When judge agreed is None, DROP the judge term and renormalize the rest — an unjudged decision must not be penalized False or faked True . parts = {"model self": s.model self, "verification": s.verification, "historical": s.historical} if s.judge agreed is not None: parts "judge" = 1.0 if s.judge agreed else 0.0 return sum WEIGHTS k v for k, v in parts.items / sum WEIGHTS k for k in parts The exact weights matter less than the principle: a single source of confidence is a single point of failure. Note model self is deliberately the smallest weight — a model that’s confidently wrong gets dragged down by failed verification or a disagreeing judge, so “confidently wrong” can’t on its own clear the bar. The three other signals are also the ones you can check : deterministic verification either passed or didn’t, the judge either agreed or didn’t, and historical accuracy is a measured fact about this slice. The model’s opinion of itself is the one input you can’t independently verify, which is exactly why it gets the least say. A score is only useful if it’s calibrated — if “0.9” means “right ~90% of the time.” Check it: bucket decisions by predicted confidence and measure real accuracy per bucket a reliability diagram . predicted bucket mean predicted actual accuracy verdict ---------------- -------------- --------------- --------------------------------------- 0.90 – 1.00 0.95 0.71 ⚠ overconfident — recalibrate / raise T 0.80 – 0.90 0.85 0.84 ✓ well-calibrated 0.70 – 0.80 0.75 0.77 ✓ < 0.70 0.55 0.52 ✓ correctly unsure If your top bucket is right 71% of the time, your automation threshold is too loose. The binning that produces that table and the gap to alert on : python def reliability samples, bins=10 : samples: list confidence, was correct: 0|1 rows, ece, n = , 0.0, len samples for b in range bins : lo, hi = b/bins, b+1 /bins last = b == bins - 1 bucket = s for s in samples if lo <= s 0 < hi or last and s 0 == 1.0 top bin includes 1.0 if not bucket: continue conf = sum c for c, in bucket / len bucket acc = sum ok for , ok in bucket / len bucket rows.append round lo, 1 , round conf, 3 , round acc, 3 , len bucket ece += len bucket / n abs conf - acc return rows, ece A weighted sum is not calibrated by construction — compose returns a number in 0,1 , not a probability. Close the loop: fit a monotonic map from raw composed score → empirical accuracy and route on that . python from sklearn.isotonic import IsotonicRegression calibrate = IsotonicRegression out of bounds="clip" .fit raw scores, correct p correct = calibrate.predict composed 0 THIS is what the threshold compares against Recompute on a rolling window — calibration drifts with the model and inputs. Prefer Brier score or the signed per-bucket gap over ECE for alerting ECE is binning-sensitive and can read ~0 for a miscalibrated model . And mind the independence caveat : weighting historical slice-accuracy into compose and then calibrating per slice double-counts — encode slice reliability in one place, not both. Once composed and calibrated, automation is a threshold — and the right T is per slice , set from the calibration data: php def route conf: float, slice key: str, thresholds: dict str, float - str: T = thresholds.get slice key, 0.95 default conservative if conf < ABSTAIN FLOOR: return "abstain" too unsure to even recommend return "auto" if conf = T else "hitl recommended" Routing here is a slimmed view of Level 1’s four-way vocabulary: abstain is a new floor below the human-review states for inputs too uncertain to even pre-fill a draft , while Level 1’s reject failed verification and hitl required very low confidence still apply. Start T high almost everything to a human , then lower it for a slice only once its calibration proves the band is safe. You can defend every automated band by pointing at the accuracy data that justified it. The highest form of this is abstention — declining to decide. An agent that says “I’m not confident; a human should look” is more trustworthy than one that always answers. Make it a first-class outcome, not a failure path. It’s also your best defense against unknown unknowns: you can’t enumerate every weird input production will send, but a calibrated score plus an abstain floor means the weird ones fall through to a human automatically. ABSTAIN FLOOR = 0.40 below this: don't even propose; escalate as "needs human" The whole system tunes around three levers, and calibration data tells you which way to move each: lever raise it lower it ------------------- ------------------------------------------ --------------------------------------- threshold T fewer auto-errors, more human load more automation, more risk model self weight trusts the model more risky leans on verification/judge abstain floor fewer bad auto-proposals, more escalations fewer escalations, more noise to humans A few anti-patterns reliably wreck this layer: using raw model confidence as the dial miscalibrated — compose instead ; a single global threshold different slices have wildly different accuracy ; never re-checking calibration it rots — a quarter-old threshold is a quarter-old risk model ; and treating abstention as an error it’s the system correctly recognizing its competence edge — measure it, don’t suppress it . One of the signals in compose was j udge agreed . Here’s where it comes from — and it earns its own section, because for high-stakes decisions it’s the difference between a confident guess and a checked one. A single model pass is a single point of failure. It’s confident when it’s wrong, it has characteristic blind spots, and you can’t tell a good answer from a plausible one by looking at the output. The cheap, effective insurance is a judge : a second, independent model that evaluates the first’s decision before you act on it. @dataclass class Verdict: agrees: bool confidence: float failure mode: str | None if it disagrees, why def judged decision inputs, primary out, judge - Verdict: return judge.evaluate inputs=inputs, proposed=primary out narrow question: is this right? Three design choices make it work. Make the judge independent — and skeptical. If the judge is the same model with the same prompt, it shares the blind spots and rubber-stamps. Use a different model family where you can, a different framing, and prompt it to refute , not confirm. You are a strict reviewer. You will be given INPUTS and a PROPOSED DECISION made by another system. Your job is to find the strongest reason the PROPOSED DECISION is WRONG, unsafe, or unsupported by the inputs. Do not be agreeable. If, after genuinely trying to refute it, you cannot, then agree. Return JSON: {"agrees": bool, "confidence": 0..1, "failure mode": string|null} INPUTS: {{inputs}} PROPOSED DECISION: {{proposed}} A judge told to refute catches far more than one told to approve — the framing does real work here, because an agreeable reviewer asked “is this right?” will find a way to say yes, while an adversarial one asked “where is this wrong?” surfaces the failure mode you’d otherwise only discover in production. Sample — don’t judge everything. Judging doubles model cost. Spend it where stakes or uncertainty are high: decision class judge policy -------------------------------------- ------------------------------------------------------ highest-impact / irreversible 100% judged borderline confidence near threshold always judged everything else random sample, e.g. 5–10% a continuous quality probe php def should judge d, thresholds, rate=0.08 - bool: if d.stakes == "high": return True if abs d.confidence - thresholds.get d.slice, 0.95 < 0.05: return True borderline per-slice T, safe default return deterministic hash d.decision id % 10000 < rate 10000 deterministic + sub-1% safe Treat disagreement as a routing signal — and a metric. When the judge disagrees, that decision goes to a human, every time. And the rate of disagreement is one of your best health signals — a spike means an attack, a bad deploy, or a model regression. php def combine primary out, verdict: Verdict - tuple: emit "judge invocations total" if not verdict.agrees and verdict.confidence = 0.6: confident refutation → human, every time emit "judge disagreements total" return route to human primary out, reason=verdict.failure mode , False judge agreed = False if not verdict.agrees: weak/low-confidence refutation: INCONCLUSIVE return primary out, None don't fake agreement — None makes compose drop the judge term return primary out, True genuine agreement → judge agreed = True for compose One gotcha will bite you: nondeterminism in tests. If the judge samples randomly and that path runs in your test suite, your tests become flaky — the same input judges on one run, not the next, so assertions intermittently fail. It looks like a mysterious bug; it’s your sampling rate leaking into a deterministic test. Pin the rate in tests, and use a deterministic hash of the decision id not RNG so even sampled behavior is reproducible: production samples; tests pin the rate. Any randomness in an agent must be injectable. judge = Judge model=secondary, sampling rate=0.0 if TESTING else 0.08 You’re not making the model perfect; you’re building a system more reliable than any single call — on a judged decision, two independent models must agree before you act, and their disagreement is the system raising its hand. For the decisions that matter, it’s a remarkably cheap way to buy a lot of safety. The dial and the judge both end at the same place: a decision routed to a person. But “route to a human” is where a lot of otherwise-good systems quietly fail — not in the model, in the handoff . Done well, the human is a genuine safety layer and a source of training signal. Done badly, you’ve built an “Approve” button people click without reading — worse than no human, because it manufactures false accountability. The lazy handoff shows the proposed answer and Approve/Reject. Under volume, humans approve — the proposal anchors them, rejecting takes effort, the queue is long. Now you have a human in the loop who isn’t checking anything, plus an audit trail that falsely says “a person reviewed this.” Good HITL UX makes a real judgment cheap and points attention where it matters. Show a draft with evidence, not a verdict to bless. The single biggest lever against rubber-stamping is showing the reasoning and the evidence — a human can’t check a conclusion they can’t see the basis for. { "proposed": { "decision": "approve", "fields": { "amount": 250 } }, // editable draft, not a verdict "confidence": 0.62, "routed because": "confidence below threshold 0.62 < 0.85 ", // tell them where to look "reasoning": "matched policy 4.2", "no prior flags" , // the WHY, not just the what "evidence": { "label": "policy", "ref": "...", "snippet": "..." } , "stakes": "reversible | high cost | irreversible", // calibrate their attention "alternatives": { "decision": "reject", "would trigger": "..." } } Capture every action as signal. Each human action is a label — record it, structured, against the decision. An edit is the richest signal: a corrected answer, gold for your golden set and for spotting where the agent is systematically off. python def on review decision id, review : ledger.append outcome decision id, review audit: who decided, what, when emit "human override rate", 1.0 if review "action" = "approve" else 0.0, slice=slice of decision id if review "action" == "edit": golden set.add candidate decision id, review "corrected" edits become eval cases Close the loop back to the dial. HITL and the confidence threshold are one feedback loop, and this is what makes the whole layer self-improving rather than static: observation on a slice action ---------------------------- -------------------------------------------------------------- humans approve ~unanimously candidate to automate — lower T for that slice frequent overrides raise T or pull the slice back to human-only many edits of the same field the agent is systematically wrong there — fix the prompt/logic You can open up automation defensibly — point at the approval rate that justified each band. This is the loop closing on itself: the judge’s disagreements and the humans’ edits both flow back into the historical and slice-accuracy signals that compose reads, so every routed decision makes the next batch of automatic ones a little better calibrated. And don’t drown the humans. Routing everything to a person isn’t safety; it’s a DoS on your reviewers, and a flooded queue gets rubber-stamped. The point of good confidence and routing is that only genuinely uncertain decisions reach a person — few enough to get real attention. If the queue is overwhelming, fix it upstream better confidence, automate the proven-safe slices , not with more reviewers. Two more traps worth naming: making reject harder than approve one click vs a five-field form biases toward approval , and omittin g routed because the human doesn't know where to look, so they skim . Confidence is the linchpin of safe automation, and it’s one mechanism in three parts. Compose the number from independent signals — never the model’s own claim — and calibrate it against real outcomes so “0.9” actually means 90%. For the decisions that matter, have an independent, skeptical judge try to prove the first model wrong, sampled by stakes and kept deterministic so your tests stay sane; its agreement feeds the score and its disagreement raises a hand. Everything below the line goes to a human , handed off with the draft, the reasoning, and why it was routed — disagreeing as cheap as approving, every action captured as signal that tunes the thresholds back down or up. A system built this way knows when it doesn’t know. It can be trusted with more, because the cases it shouldn’t handle route themselves away — and the appearance of oversight, the most dangerous outcome of all, is exactly what it refuses to manufacture. Series: Running LLM systems in production — Level 3 of 6: Confidence.