Building Multi-Agent AI Systems: A Developer's Primer A developer primer outlines how splitting work across multiple specialized AI agents with a coordinator — agent orchestration — avoids the context-window bloat, role confusion, and opaque debugging that degrade single agents as tasks grow. It cites Anthropic's finding that a Claude Opus 4 lead agent with Claude Sonnet 4 subagents outperformed single-agent Claude Opus 4 by 90.2% on an internal research eval while consuming roughly 15× more tokens, and recommends schema validation, payload versioning, and structured output at every agent boundary to prevent silent schema drift. Imagine an AI agent with 15 tools in a long conversation. Initially, it works fine. But as the conversation grows, it starts skipping steps. Then it tries to use a tool that doesn't exist, with made-up inputs. Nothing crashes. The agent just slowly gets worse at its job. That's why more developers now build multi-agent systems https://databyteworks.com/blog/multi-agent-ai-systems-explained-how-ai-agent-development-transforms-business-workflows/ . Instead of one agent doing everything, you split the work across several smaller agents, each with a single, clear job. A coordinator decides which agent handles which task. Developers call this agent orchestration . In this primer, you'll learn the main design patterns, how agents pass information to each other, how to build a simple orchestrator, and how to keep the whole system safe in production. A single agent works well for narrow jobs. It breaks in three predictable ways as the job grows. Context window bloat Every tool schema, retrieved document, and prior turn competes for the same context window. Add tools and history, and the most important instructions get diluted. The model doesn't fail loudly; it starts missing steps. One model doing three jobs You ask the same prompt to plan, retrieve, and execute. Each job wants different instructions, different tools, and often a different model. Cramming them together forces compromises in all three. Opaque debugging When it goes wrong, you get one long trace. There are no natural seams, so you can't tell whether the plan, the retrieval, or the tool call broke. Splitting the work into agents gives you those seams for free: each agent has its own input, output, and logs. The trade-off is costly. Anthropic https://www.anthropic.com/engineering/multi-agent-research-system found that its multi-agent research system, a Claude Opus 4 lead agent with Claude Sonnet 4 subagents, outperformed single-agent Claude Opus 4 by 90.2% on its internal research eval. The same system used about 15× more tokens than a chat interaction. Multi-agent wins on hard, parallelizable work. On simple tasks, you pay more for no gain. Three Agent Orchestration Patterns, and Where Each One Breaks Most production multi-agent systems use one of three topologies. Pick based on where you can afford to fail. Start with the orchestrator-worker. It gives you one place to log, route, and stop things, and you can move to a hybrid later. Most teams who start peer-to-peer end up adding a coordinator anyway, just to see what's happening. This is where most multi-agent prototypes quietly fail. You have two ways to move state between agents, and each breaks differently. Shared state feels simpler at two agents. At five, nobody can say which agent wrote a value, or when. Use a shared store for read-mostly reference data, such as product catalogs or embeddings. Pass working state as explicit payloads. Schema drift: the silent killer Agent A stops sending customer tier because someone simplified its prompt. Agent B still expects it, defaults to None, and prices every customer at the standard rate. No error fires. Revenue just leaks. Three habits prevent this: Validate at every boundary . Parse each agent's output into a schema Pydantic, Zod or JSON Schema and fail loudly on missing fields. Version the payload . Add a schema version field so a consumer can reject a shape it doesn't understand. Never let free text cross a boundary . Ask the model for structured output, and treat prose as a bug in an agent's contract. from pydantic import BaseModel class PricingRequest BaseModel : schema version: int = 2 customer id: str customer tier: str required: B fails loudly if A drops it seats: int A Minimal Orchestrator Pattern Here's an orchestrator–worker skeleton in plain Python, standard library only. It runs as-is. Swap the lambdas for LLM calls and classify for a model that returns one label from a fixed set. It does three things most tutorials skip: every agent returns a structured result, every agent has its own timeout and circuit breaker, and every call writes an audit line. from concurrent.futures import ThreadPoolExecutor from dataclasses import dataclass, field import time @dataclass class AgentResult: agent: str status: str "ok" | "error" | "skipped" data: dict = field default factory=dict class CircuitBreaker: def init self, max failures=3, cooldown s=60 : self.max failures, self.cooldown s = max failures, cooldown s self.failures, self.opened at = 0, None python def allow self : if self.opened at and time.time - self.opened at < self.cooldown s: return False return True def record self, ok : self.failures = 0 if ok else self.failures + 1 self.opened at = time.time if self.failures = self.max failures else None class Agent: def init self, name, handler, timeout s=10 : self.name, self.handler, self.timeout s = name, handler, timeout s self.breaker = CircuitBreaker pool = ThreadPoolExecutor max workers=4 def run agent agent, request : if not agent.breaker.allow : return AgentResult agent.name, "skipped", {"reason": "circuit open"} try: data = pool.submit agent.handler, request .result timeout=agent.timeout s agent.breaker.record ok=True return AgentResult agent.name, "ok", data except Exception as exc: includes timeouts agent.breaker.record ok=False return AgentResult agent.name, "error", {"error": repr exc } def classify request : Swap in an LLM call that returns one label from a fixed set. text = request "text" .lower if "price" in text or "quote" in text: return "pricing" if "book" in text or "meeting" in text: return "scheduling" return "fallback" AGENTS = { "pricing": Agent "pricing", lambda r: {"quote usd": 1200} , "scheduling": Agent "scheduling", lambda r: {"slot": "2026-10-08T10:00"} , "fallback": Agent "fallback", lambda r: {"route to": "human"} , } def orchestrator request : task = classify request result = run agent AGENTS task , request if result.status = "ok": result = run agent AGENTS "fallback" , request print f"audit agent={result.agent} status={result.status} data={result.data}" return result if name == " main ": orchestrator {"text": "Can I get a price for 3 seats?"} orchestrator {"text": "Book a meeting next week"} orchestrator {"text": "My login is broken"} Output: audit agent=pricing status=ok data={'quote usd': 1200} audit agent=scheduling status=ok data={'slot': '2026-10-08T10:00'} audit agent=fallback status=ok data={'route to': 'human'} Note what the orchestrator doesn't do: it never parses free text from a worker. Each agent returns an AgentResult, so the orchestrator, and any agent downstream, can act on it reliably. In production, swap print for structured logging and the in-memory breaker for a shared one Redis works so every replica sees the same state. The biggest production risk isn't one agent failing. It's one agent failing quietly. Cascading failures An extraction agent misreads a date. The validation agent trusts it. The scheduling agent books the wrong slot, and the notification agent confirms it to the customer. Every agent did its job correctly with bad input. In a multi-agent system, errors compound at every handoff instead of staying put. Containment is an architecture problem A 2026 Dataiku survey https://gulfnews.com/technology/uae-firms-lead-the-world-in-deploying-ai-agents-but-most-cant-contain-a-rogue-one-quickly-survey-finds-1.500692105 , run by Harris Poll, found that only 5% of UAE CIOs said their organisation could reliably identify and contain a problematic AI agent within one to two hours. The global figure was 10%. Read that as an engineer, and it's rarely a process failure. It's an AI architecture gap: no circuit breaker at the orchestration layer, no per-agent kill switch, and no way to halt one agent without taking down the whole pipeline. If stopping one misbehaving agent means a redeploy, containment will always take hours. The cost of skipping this shows up later. Gartner https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027 predicts that over 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, or inadequate risk controls. Weak controls are the one cause on that list that engineers directly own. The takeaway: design the circuit breaker before you need it, not after your first incident. Each of these is cheap on day one and painful to retrofit after launch. A few specifics: Time out per agent, not just per request - A global 30-second timeout tells you something was slow. A per-agent timeout tells you which agent, and lets the orchestrator fall back while the others finish. Make the breaker controllable - Expose a flag your on-call engineer can flip to disable one agent without a redeploy. That flag is your containment plan. Log the decision, not just the call - Record which agent ran, its input, its output, and what the orchestrator did next. One end-to-end trace won't show you where a cascade started. Gate irreversible actions - Let agents draft the payment, the email, or the filing. Let a human, or at least a deterministic rule, approve it. Picking a Framework The framework matters less than the patterns above, but it shapes how fast you move. Here's the landscape most teams compare today, with no endorsement. One note if you're evaluating AutoGen : Microsoft now describes Agent Framework as its successor and keeps AutoGen on bug fixes and security patches. For new builds, start with Agent Framework. Don't underrate rolling your own. If your workflow has three agents and a clear routing rule, a framework can add more abstraction than value. Good LLM development https://databyteworks.com/llm-development/ practice applies either way: pin model versions, test each agent in isolation, and keep prompts in version control next to the code. Start with the orchestrator–worker. Pass typed payloads instead of shared state. Give every agent its own timeout, breaker, and audit log. Then scale the number of agents, not the size of their prompts. If you're scoping a production multi-agent build, not just a prototype, here's how we approach the architecture and integration side: AI Agent Development https://databyteworks.com/ai-agent-development/ . FAQs - How many agents should a first multi-agent system have? As few as the workflow needs, usually two to four. Every agent adds a handoff, a schema, and a failure point. Split an agent only when its prompt, tools, or context window clearly overload it. - Do all agents need to run on the same LLM? No, and they often shouldn't. Use a stronger model for the orchestrator or planner and smaller, cheaper models for narrow workers such as classification or extraction. Anthropic's research system pairs an Opus lead agent with Sonnet subagents for this reason. - How do I test a multi-agent system before production? Test each agent in isolation against fixed inputs and expected structured outputs, like a unit test. Then replay logged real requests through the full pipeline and compare results. Your audit logs from the orchestrator double as the replay dataset. - Should agents share one vector database? Share it for read-mostly reference data, such as documentation or a product catalogue. Don't use it as a scratchpad for working state between agents. Pass that state as typed payloads so you can trace who changed what. - How do I stop multi-agent costs from spiralling? Measure tokens per agent, not just per request. Cap each agent's iterations, cache repeated retrievals, and route simple requests past the expensive agents entirely. Multi-agent systems can use around 15× the tokens of a chat, so budget per agent. - When should I use a framework instead of my own orchestrator? Pick a framework when you need durable checkpointing, long-running workflows, or complex branching and loops. Roll your own when you have a handful of agents and a clear routing rule. You can migrate later if you keep agents behind a clean interface.