Agentic Systems in Production: Patterns That Survive Real Traffic Production agentic systems require deterministic orchestrators wrapping non-deterministic reasoners, with idempotent tools, hard budget caps, and human-in-the-loop gates on irreversible actions to survive real traffic, as single-pass LLM calls fail due to orchestration, identity, and observability issues rather than model failures. The Problem Single-pass LLM calls don't survive contact with production. The moment you give a model tools that mutate state — booking flights, processing refunds, opening pull requests, rerouting shipments — every property you took for granted in a stateless API breaks: retries are no longer idempotent, latency is unbounded, the action space is non-deterministic, and the failure mode is now "wrong action executed" rather than "wrong text returned" Source 2 source-2 Source 16 source-16 . Most production agent failures aren't model failures; they're orchestration, identity, and observability failures dressed up as model failures Source 17 source-17 . The Shape The pattern that holds up: a deterministic orchestrator wrapping a non-deterministic reasoner, with idempotent tools, hard budget caps, and a human-in-the-loop gate on irreversible actions Source 5 source-5 Source 21 source-21 . Copy-paste skeleton: python import asyncio, time, uuid, logging from dataclasses import dataclass, field log = logging.getLogger "agent" @dataclass class RunBudget: max steps: int = 12 max tokens: int = 100 000 max usd: float = 2.00 deadline s: float = 90.0 tokens used: int = 0 usd used: float = 0.0 steps: int = 0 started: float = field default factory=time.monotonic def check self : if self.steps = self.max steps: raise BudgetExceeded "steps" if self.tokens used = self.max tokens: raise BudgetExceeded "tokens" if self.usd used = self.max usd: raise BudgetExceeded "usd" if time.monotonic - self.started self.deadline s: raise BudgetExceeded "deadline" class BudgetExceeded Exception : pass class CircuitOpen Exception : pass TOOL ALLOWLIST = {"search kb", "get order", "draft refund"} HITL REQUIRED = {"issue refund", "send email", "create ticket"} class CircuitBreaker: def init self, threshold=5, cooldown=30 : self.fail = 0; self.threshold = threshold self.opened at = 0; self.cooldown = cooldown def allow self : if self.fail < self.threshold: return True if time.monotonic - self.opened at self.cooldown: self.fail = self.threshold - 1 return True return False def record self, ok : if ok: self.fail = 0 else: self.fail += 1 if self.fail == self.threshold: self.opened at = time.monotonic BREAKERS = {} async def call tool name, args, idempotency key, breaker : if name not in TOOL ALLOWLIST: return {"error": f"tool '{name}' not allowlisted"} if not breaker.allow : raise CircuitOpen name for attempt in range 3 : try: res = await asyncio.wait for TOOLS name args, idempotency key=idempotency key , timeout=5.0, breaker.record True return res except asyncio.TimeoutError, TransientError : await asyncio.sleep 2 attempt + attempt 0.1 breaker.record False return {"error": "tool failed after retries"} async def hitl gate action, args, run id : approval = await approvals.request run id=run id, action=action, args=args, ttl s=600 return approval.decision == "approve" async def run agent user msg, principal, budget=None : budget = budget or RunBudget run id = str uuid.uuid4 trace = state = {"messages": {"role": "user", "content": user msg} } while True: budget.check ; budget.steps += 1 step = await llm.plan state, tools=list TOOL ALLOWLIST | HITL REQUIRED , principal=principal, budget.tokens used += step.usage.total tokens budget.usd used += step.usage.cost usd trace.append {"run": run id, "step": budget.steps, "thought": step.thought, "action": step.action, "args": step.args} if step.action == "final": log.info "agent.done", extra={"run": run id, "steps": budget.steps} return step.answer, trace breaker = BREAKERS.setdefault step.action, CircuitBreaker idem key = f"{run id}:{budget.steps}:{step.action}" if step.action in HITL REQUIRED: if not await hitl gate step.action, step.args, run id : state "messages" .append {"role": "tool", "name": step.action, "content": "denied by human"} continue try: result = await call tool step.action, step.args, idem key, breaker except BudgetExceeded, CircuitOpen as e: state "messages" .append {"role": "tool", "name": step.action, "content": f"halt:{e}"} return await llm.summarize halt state, reason=str e , trace state "messages" .append {"role": "tool", "name": step.action, "content": result} Every step is traced, every tool call is keyed for idempotent retry, every action that mutates the world either fails closed or requires human approval, and the loop cannot exceed its step, token, USD, or wall-clock budget Source 5 source-5 Source 8 source-8 Source 26 source-26 . How It Works The agent loop itself is the ReAct pattern — observe, reason, act, repeat — wrapped around a model whose action space is constrained to a tool allowlist, with each tool described by a JSON schema the model uses for routing and parameter generation Source 13 source-13 Source 23 source-23 . The orchestrator, not the model, owns control flow: it counts steps, charges the budget, fans out to tools, and decides when to hand off to a human. "Separating the brain from the hands" — the model classifies and extracts, deterministic code applies the patch — is what keeps a hallucinated argument from becoming a hallucinated refund Source 15 source-15 . Idempotency is the load-bearing property. Tool calls to external APIs fail transiently; retry with exponential backoff is mandatory, but only safe when the tool checks for an existing record with the same idempotency key before creating a new one Source 5 source-5 Source 8 source-8 . The circuit breaker — closed, open, half-open — is the same Hystrix pattern Netflix taught the industry; in an agent context it stops a degraded downstream from burning the entire token budget on doomed retries Source 19 source-19 Source 7 source-7 . Bulkhead the breakers per-tool so a flaky email API doesn't poison the search path. Identity and authorization are the part most demos skip. Agentic context is autonomous, dynamic, multi-system; the user's identity must propagate through the orchestrator, sub-agents, and MCP servers to whatever resource finally executes the write, or you create a confused-deputy problem at scale Source 2 source-2 Source 33 source-33 . Each agent should have a unique identity, least-privilege scoped to its task, with just-in-time provisioning for sensitive credentials and a narrow tool catalog so a compromised sub-agent has nowhere to pivot Source 12 source-12 Source 12 source-12 Source 16 source-16 . Prompt injection through retrieved content is real — five poisoned documents can flip behavior with 90% success in published research — so the orchestration layer must validate tool args, not trust the model's claim about them Source 16 source-16 . The observability layer is non-negotiable. Catchpoint's framing — "what the AI decided / what it executed / where it broke" — is the right schema for traces, because page-load and API-latency dashboards don't tell you whether intent was actually fulfilled Source 17 source-17 Source 17 source-17 . Distributed trace IDs link the LLM call to every tool invocation; cost-per-task and steps-per-task are the leading indicators of orchestration regressions long before user-facing errors appear Source 8 source-8 . user ──▶ orchestrator ──▶ planner LLM │ │ thought + action │ budget/step ◀────┘ │ ├──▶ allowlist check ──▶ HITL gate if mutating │ │ approve/deny ├──▶ circuit breaker ──▶ tool idempotent, timeout, retry │ │ result │ trace + cost ◀───────────────┘ ▼ audit log / observability When It Breaks | Condition | What happens | Use instead | |---|---|---| | Single mega-tool wraps a 40-parameter API | enum -constrained targets; resolve IDs server-side from natural language Source 15 source-15 Source 26 source-26 Source 1 source-1 Source 1 source-1 Source 1 source-1 Source 20 source-20 Source 10 source-10 Source 11 source-11 Source 7 source-7 Source 19 source-19 Source 9 source-9 Source 21 source-21 Source 26 source-26 Source 32 source-32 Source 4 source-4 Source 29 source-29 Source 29 source-29 Source 30 source-30 Source 6 source-6 Source 18 source-18 Source 3 source-3 Source 28 source-28 Source 22 source-22 Source 24 source-24 Source 22 source-22 Source 3 source-3 Source 25 source-25 Source 3 source-3 Source 27 source-27 Source 31 source-31 Source 27 source-27 CEMENT Brick If you ship an agentic workflow without budget caps, idempotent tools, a deterministic orchestrator, propagated identity, and a HITL gate on irreversible actions, then your first real-traffic incident will be unrecoverable, because the same autonomy and non-determinism that make agents useful turn every missing guardrail into a load-bearing failure mode — and unlike a stateless API, you cannot roll back the actions an agent has already taken in the world Source 9 source-9 Source 21 source-21 Source 17 source-17 . Sources Build, Reuse, or Hybrid? How Orchestration Powers Agentic AI https://www.youtube.com/watch?v=tNQPNBQC5kg How to Pass Context in an Agentic AI Flow https://www.youtube.com/watch?v=UC4vDpSJCkM - Engineering Docs How AI Agents and Decision Agents Combine Rules & ML in Automation https://www.youtube.com/watch?v=-mldKsBR0UM - Engineering Docs Enhancing AI Agents Through Fine Tuning & Model Customization https://www.youtube.com/watch?v=aQuCTWhiiPg - Engineering Docs - Engineering Docs Risks of Agentic AI: What You Need to Know About Autonomous AI https://www.youtube.com/watch?v=v07Y4fmSi6Y behind-the-streams-real-time-recommendations-for-live-events-e027cb313f8f https://netflixtechblog.com/behind-the-streams-real-time-recommendations-for-live-events-e027cb313f8f Behind the Streams: Real-Time Recommendations for Live Events Part 3 https://netflixtechblog.com/behind-the-streams-real-time-recommendations-for-live-events-e027cb313f8f?source=rss----2615bd06b42e---4 What Are AI Identities? Understanding Agentic Systems & Governance https://www.youtube.com/watch?v=AuV62XbiZcw - Engineering Docs Building Tools for AI Agents https://www.youtube.com/watch?v=ov-HUEVrgOk - Engineering Docs - Engineering Docs How to Monitor AI Agents in Commerce Systems https://www.catchpoint.com/blog/how-to-monitor-ai-agents-in-commerce-systems AI Dev 25 x NYC Nicholas Clegg: How AWS Moved Beyond Orchestration with Strands SDK https://www.youtube.com/watch?v=lVgrowsPASU - Engineering Docs AI agents in action: From pilots to outcomes at scale https://www.youtube.com/watch?v=v-Q0hyKl88I Why AI Agents Need A Human in the Loop Now https://www.youtube.com/watch?v=cmEJ-5zYKHA Uber: Leading engineering through an agentic shift - The Pragmatic Summit https://www.youtube.com/watch?v=i1tZN41VKcE - Engineering Docs LLM vs. SLM vs. FM: Choosing the Right AI Model https://www.youtube.com/watch?v=AVQzG2MY858 - Engineering Docs - Engineering Docs The future of agentic coding: conductors to orchestrators https://addyosmani.com/blog/future-agentic-coding/ Orchestrator Agents & MCP: How AI Agents Drive Automation https://www.youtube.com/watch?v=Ons1Fv3IE4U Building Decision Agents with LLMs & Machine Learning Models https://www.youtube.com/watch?v=mRkJTXDromw Designing AI Decision Agents with DMN, Machine Learning & Analytics https://www.youtube.com/watch?v=Wtpwva8t1vs Your AI coding agents need a manager https://addyosmani.com/blog/coding-agents-manager/ Building an AI Agent Governance Framework: 5 Essential Pillars https://www.youtube.com/watch?v=5hK7pQsvpy0 Securing Agentic Frameworks https://www.youtube.com/watch?v=MLPMpE4wJTQ