I once watched a group chat of eleven agents spend forty minutes debating whether a research summary was "comprehensive enough" while the actual user question sat unanswered in the scrollback. Every agent was working. Every agent was contributing. The conversation was rich, thoughtful, and completely useless. By the time the moderator synthesized a final answer, the user had closed the tab.
That was the day I understood the trap: group chat orchestration works beautifully until it doesn't, and the "doesn't" arrives with no warning and no error message.
The group chat pattern is the most natural multi-agent topology to reach for. Agents share a conversation. A moderator picks who speaks next. Each agent sees the full history and contributes based on its role. AutoGen's GroupChatManager, CrewAI's conversational crews, and a dozen frameworks implement some version of this. Microsoft's own orchestration training materials describe it as "agents collaborating in a shared conversation," and it's the pattern most teams build first because it mirrors how human teams actually work.
It works well for three to five agents on a bounded task. The research is blunt about what happens next. When you scale from three agents to twelve specialized agents, flat patterns exhibit three failure modes. Coordination overhead grows quadratically—with N agents, you potentially manage N(N-1)/2 relationships. A group chat with ten participants becomes chaotic when agents speak out of turn or contradict each other.
I hit all three in the same week. My moderator started routing work to agents that had already contributed, because the conversation history was too long for it to track who had done what. Two agents produced contradictory conclusions and the moderator, lacking a protocol for adjudication, averaged them into a mush. And the token bill arrived at the end of the month with a number that made me close my laptop and walk outside.
The cost model is the part that should be on every architecture slide.
In a group chat, every agent reads the full prior conversation each turn. The token cost grows roughly quadratically with agents × rounds. With N agents and M rounds, total tokens scale roughly as N × M × (N × average turn length), because every agent re-ingests every earlier turn on every turn it takes.
The empirical numbers are worse than the formula suggests. Anthropic's internal analysis found that agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats. A 5-agent crew doesn't cost 5× a solo agent on input tokens—it costs closer to 6×, and once you stack memory, delegation, and verbose mode it climbs to 8–15×.
AutoGen GroupChat and LangGraph StateGraph re-send the full append-only history each step, so the shared-prefix prefill grows super-linearly absent reuse. The measured cross-framework multipliers: CrewAI 4.93×, AutoGen/LangGraph 7.82×.
I had been blaming the model. The model was fine. The architecture was re-reading the entire conversation on every turn and calling it collaboration.
The COSMIC paper from IEEE's 2026 conference catalogued the specific failure modes that emerge as group chats scale. Each of these patterns suffers from concrete shortcomings that matter in high-dependency tasks: unstable turn-taking and termination, fragmented shared context, and weak global oversight. In group-chat settings, agents communicate via free-form natural language, which encourages flexibility but also leads to over-communication, redundant messages, and difficulty deciding when to stop.
I recognized every one of those in my traces.
Unstable turn-taking. Without a strict speaker-selection policy, agents talk over each other. One agent asks a clarifying question, another agent answers a different question, a third agent responds to neither and repeats work the first agent already did. The conversation moves forward but progress doesn't.
Fragmented shared context. Each agent maintains its own view of the conversation, and those views diverge. The moderator's summary at the end reflects whatever it happened to retain, not what actually happened. The 2026 MAST taxonomy found that inter-agent misalignment accounts for 36.9% of failures, with coordination breakdowns—communication failures, state synchronization errors, conflicting objectives—accounting for a plurality of documented failures.
Weak global oversight. No single agent has the full picture. The moderator sees the conversation but doesn't see what each agent actually did versus what it said it did. The GitHub issue from the LangChain planner-executor split documents the exact symptom: the planner calls web_fetch three times, writes the answer into its handoff text, and the executor—receiving only the text—calls web_fetch three more times, walks the same wrong paths, and burns the same latency again.
Broadcast storms. The openwalrus analysis documents the extreme case: with N agents, a single user message generates N responses, each of which generates N-1 follow-up responses. Token consumption grows exponentially. A documented case of a circular agent relay persisted for 9+ days, consuming 60,000+ tokens.
The pattern is consistent across frameworks and deployments.
Microsoft's GroupChatOrchestration had a documented token explosion bug when MaximumInvocationCount > 3. The root cause was that RoundRobinGroupChatManager wasn't usable as-is with unbounded rounds.
AutoGen's conversational overhead enables superior adaptation—0.35 Plan Adaptation BCS, 28% better than CrewAI—but at 4.1M tokens, a 27% overhead compared to graph-based approaches. You pay for flexibility in tokens.
The one production retrospective I found where group chats were used organically reported that nobody used them despite 100+ cross-agent messages. The root cause: the lead's hub-and-spoke routing created high activation energy for manual group creation. Zero of six agents used group chats organically. The one case where groups would have helped—five developers editing the same file in a serialized queue—was handled by a queue instead.
The group chat was the right topology for the problem. The routing layer made it too expensive to use.
The fix isn't abandoning group chat. It's adding a routing layer that decides when the group conversation is the right medium and when a direct handoff or a parallel fan-out is better.
The StigmergyRouter paper describes exactly this: a fault-aware routing layer that maintains clustered pheromone memory over semantic query embeddings. The paper reports both the pure mechanism and a HybridStigmergyRouter that combines semantic priors with heartbeat-driven updates. The routing layer doesn't replace the group chat. It decides which messages need to go to the whole group, which need to go to a specific specialist, and which need to go nowhere because the work is already done.
The Maestro group chat system implements the hub-and-spoke version: a central moderator coordinates messages, decides which agents to delegate to via @mentions, and the router auto-adds any mentioned agents not yet in the chat. This is the practical middle ground: the group chat exists for the moments when broad collaboration is needed, but the router prevents every message from becoming a broadcast.
The AgenticX Group Chat and Routing system supports a hybrid "Workforce" bridge that transforms complex tasks into structured multi-agent orchestrations—round-robin, and team-based routing. The routing layer is what makes the group chat affordable.
I rebuilt my system around three changes.
A semantic router in front of the group chat. Every incoming message hits a lightweight classifier—the Auroic Router 0.6B runs on CPU in under 4 seconds and outputs a structured routing decision from a 5-message history window and up to 3 unprocessed candidate messages. The router decides whether the message needs the full group, a specific specialist, or a direct handoff to the user. Most messages never enter the group conversation.
Turn-taking with a budget. When a message does enter the group, the moderator enforces a hop budget. Microsoft's own guidance recommends a chain depth limit of three to four handoffs before escalation. My moderator has a max_group_turns parameter. When the budget is exhausted, it either produces a best-effort synthesis or escalates to a human. No more forty-minute debates about whether a summary is comprehensive enough.
Context compaction at the boundary. Every agent that enters the group chat receives a compressed summary of prior turns, not the full transcript. The Sentex approach—extracting only the sentences relevant to the next agent's task—cuts the token cost per turn without losing the signal. The shared conversation becomes a summary of what happened, not a transcript of every utterance.
Use group chat when the task genuinely requires deliberation. When multiple perspectives need to be surfaced before a decision, when no single agent has the full answer, when the value is in the interaction. Use it for bounded tasks with a clear stopping condition.
Use hybrid routing when group chat is part of a larger workflow. When most messages need a specific specialist, when the group is the exception rather than the rule, when the token cost of every message going to every agent is more than you can justify.
Don't use group chat when the task is a pipeline. If stage B genuinely depends on stage A, use a sequential chain. The group chat will serialize the work anyway—and pay quadratic token costs for the privilege.
Don't use group chat when the task is parallelizable. If six independent sources need analysis, fan out and fan in. The group chat will run them sequentially and call it collaboration.
The hybrid routing layer buys you token efficiency and turn-taking discipline. It costs you the emergent flexibility that made group chat appealing in the first place.
Every routing decision is a chance to misroute. A message that should have gone to the group gets routed to a single specialist, and the specialist misses context that another agent had. The router's accuracy becomes the system's accuracy. The StigmergyRouter paper is honest about this: stigmergic memory is useful as an adaptive overlay on strong semantic routing, not as a replacement for it.
And you're accepting more infrastructure. A router, a budget, a compaction layer. Every one of those is a thing that can fail independently. But here's what I've learned from watching a group chat burn 40% of its tokens on redundant messages and another 20% on looping acknowledgments: the coordination tax is real, and the only way to pay less is to route around it.
So here's my question: When your group chat finishes a task, can you tell me how many of those messages were load-bearing—or are you still paying for a conversation that mostly repeated itself?
I'd love to hear where you've landed. A tight roster and a max-round limit, a hybrid router in front of the group, a full migration to fan-out, or a group chat you're still defending—and what finally made you look at the token bill?