Stop Over-Orchestrating AI Agents: Simplicity Wins A developer reports that routing AI agent requests through orchestration frameworks—routing layers, context managers, memory brokers and dispatch queues—can triple token consumption and add over three seconds of latency compared with a flat, direct-call pattern, citing Gartner's State of AI Agents report that simpler agent designs show better ROI in production. The account argues orchestrators are justified only for genuinely parallel workloads, merged multi-agent outputs, audit trails or centralized tool-permission enforcement, and that most teams adopt the coordination infrastructure before having workloads that need it. In 2025, the default advice for anyone building AI agents was: pick an orchestration framework, wire everything through a central coordinator, and let the middleware handle complexity. We followed that advice. Three months into a build, we had a beautifully layered system where every agent request traveled through a routing layer, a context manager, a memory broker, and a dispatch queue before reaching the model that would actually do the work. Latency was high. Token bills were higher. Debugging felt like reading a stack trace through frosted glass. A debate has been building in developer communities on Hacker News and in engineering Slack groups throughout 2025 and into 2026: are orchestration frameworks solving real problems, or are they a category of over-engineering that the industry adopted before it had enough production experience to know better? The question is worth taking seriously. According to Gartner's State of AI Agents report https://www.gartner.com/en/documents/4741121 , organizations are increasingly recognizing that overly complex AI agent architectures lead to higher operational costs and reduced performance, with simpler, more focused agent designs showing better ROI in production environments. That finding matches what we saw firsthand. The core promise of an orchestration layer is coordination: one system that routes tasks, manages state, and ensures agents don't step on each other. The cost of that promise is abstraction. Every abstraction layer in an AI pipeline adds tokens. Here is the mechanism. A user sends a request. The orchestrator receives it, formats it into a routing prompt, sends that prompt to a reasoning model to determine which sub-agent should handle the task, receives the routing decision, formats a new prompt for the target agent, sends that prompt, receives the response, formats a synthesis prompt, and finally returns an answer. Each formatting step injects system context, role definitions, and state summaries that the model needs to orient itself. None of that context is free. In a direct-call pattern, the same request goes to one model with one system prompt. The work gets done. The conversation ends. The token multiplication is not theoretical. It compounds with every hop. A three-agent pipeline with a central orchestrator can easily triple the token consumption of a direct implementation handling the same task. At low volume, that difference is invisible. In production, it becomes a line item that someone has to explain to finance. Latency follows the same pattern. Each orchestration hop is a synchronous API call. If each call takes 800 milliseconds, a four-hop pipeline adds over three seconds of pure coordination overhead before any real work begins. Users notice three seconds. Fairness requires acknowledging what orchestrators do well. They shine in genuinely parallel workloads where multiple independent agents need to run simultaneously and their outputs need to be merged. They also help when you need a single audit trail across a complex multi-step process, or when different agents require different tool permissions and you want one place to enforce access control. If you are building a system where ten agents are simultaneously researching, writing, fact-checking, and formatting a document, a coordinator that manages their outputs is doing real work. The coordination cost is justified because the alternative, managing that concurrency yourself in application code, is worse. The problem is that most teams reach for orchestration frameworks before they have workloads that require them. They build the coordination infrastructure first, then fill it with tasks that a single well-prompted model could handle directly. The framework becomes the architecture, and the architecture becomes the constraint. A direct agent pattern is not a primitive or a shortcut. It is a deliberate choice to keep the call graph flat. One model receives a well-constructed prompt, uses tools if it needs them, and returns a result. The calling application handles routing logic in code, not in a prompt sent to another model. This approach has three concrete advantages. First, debugging is straightforward: you have one input, one output, and a clear log of every tool call in between. When something breaks, you know exactly where to look. Second, iteration is faster because you are editing a prompt and a tool list, not reconfiguring a multi-agent topology. Third, the cost per task is predictable because you are not paying for coordination tokens that vary based on how the orchestrator interprets the routing task. We learned a version of this lesson the hard way during our first Stripe product creation. The API call included a recurring parameter set to null . We thought omitting the value was the same as omitting the field. It wasn't. Stripe created two prices: one correct one-time payment at $297, and one spurious monthly subscription at $297. We caught it before a customer was charged monthly for a one-time product, but it took a manual archive in the Stripe Dashboard to fix. Now our factory pipeline never includes the recurring field at all, not null , not false , just absent. The lesson applies directly to agent architecture: the absence of a thing is not the same as setting it to zero. Removing an orchestration layer entirely is different from building one and configuring it to be lightweight. Absence is cleaner. For more on where direct patterns break down in production, our post on why 24/7 AI agents fail and what actually works https://dev.to/blog/why-24-7-ai-agents-fail-and-what-actually-works covers the failure modes we've seen most often. The choice between these patterns is not ideological. It is a function of your actual workload. Here is how we think about it. Use a direct agent pattern when: the task has a single clear owner, the output of one step is the input of the next in a linear chain, you need fast iteration cycles, or your token budget is a real constraint. Most customer-facing automations, content generation pipelines, and data extraction tasks fall here. Use an orchestration layer when: you have genuinely parallel workloads where agents must run concurrently, you need centralized access control across agents with different permission sets, or you are building a system where the coordination logic itself is complex enough that encoding it in application code would be harder to maintain than a dedicated coordinator. Research pipelines, multi-source data aggregation, and what ForgeWorkflows calls agentic logic, where the system must reason about its own next action, are legitimate candidates. The honest version of this matrix is that most teams need orchestration for fewer tasks than they think. Start with the direct pattern. Add coordination infrastructure only when you hit a specific problem that the direct pattern cannot solve. Do not build the coordination layer speculatively. The over-engineering tendency has a clear origin. The first wave of AI agent frameworks arrived before anyone had meaningful production data on what these systems actually needed. Framework authors, many of them coming from distributed systems backgrounds, applied patterns that work well for microservices: centralized coordination, message queues, service registries. Those patterns solve real problems in distributed computing. They do not automatically translate to AI agent pipelines, where the bottleneck is model latency and token cost rather than network throughput or service discovery. The Gartner finding cited above reflects a correction that is now underway. Organizations that shipped complex orchestration architectures in 2023 and 2024 are measuring their production costs and finding that simpler designs outperform them. This is the normal arc of a new technology category: early adoption favors complexity because complexity signals sophistication, and production experience corrects toward simplicity because simplicity is cheaper to operate and easier to fix. The developer community debate happening right now on HN and in engineering forums is that correction happening in public. It is worth paying attention to, not because the contrarian position is always right, but because the people making the argument are the ones who built the complex systems and are now living with the results. If you want a broader view of how experienced engineers are approaching these tradeoffs in 2026, our post on how top engineers actually build AI stacks https://dev.to/blog/how-top-engineers-actually-build-ai-stacks covers the patterns we see repeated across teams that ship reliably. We'd instrument token consumption per hop before committing to any architecture. The total token cost of a multi-agent pipeline is not obvious from the design diagram. We would add logging at every model call from day one, measure the coordination overhead as a percentage of total tokens, and use that number to justify or reject the orchestration layer. If coordination tokens exceed 20% of total consumption, the architecture needs a second look. We'd treat orchestration as a late-stage addition, not a foundation. The instinct to build the coordination infrastructure first is understandable but expensive. Starting with direct patterns and adding coordination only when a specific production problem demands it would have saved us significant rework. The framework should emerge from the requirements, not precede them. We'd evaluate what ForgeWorkflows calls a modular swarm pattern earlier. Rather than a single central orchestrator managing all agents, a loosely coupled set of specialized pipelines that hand off to each other through simple API calls preserves most of the coordination benefit while keeping each component independently debuggable. We came to this architecture late. It should have been the starting point for any system where more than two agents needed to interact.