cd /news/ai-agents/stop-burning-tokens-why-multi-agent-… · home › topics › ai-agents › article
[ARTICLE · art-148090] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Stop Burning Tokens: Why Multi-Agent Orchestration Beats Single-Prompt Bloat in AI Coding

A developer argues that single-prompt bloat—packing an entire codebase and task description into one large inference call—is an expensive and brittle approach to AI-assisted coding, and that multi-agent orchestration with context isolation is the better architecture. The writeup explains that attention cost scales with sequence length and that critical context buried mid-prompt is under-utilized, causing hallucinations and retries, while an orchestrator that decomposes tasks into a DAG and routes only relevant context to specialized agents can cut per-agent prompts from roughly 50,000 tokens to about 2,000.

by read8 min views1 publishedOct 9, 2026

Originally published on tamiz.pro.

When building AI-assisted development workflows, the most common failure mode is not a lack of model capability, but rather the architecture of how context is managed. Many developers default to a "single-prompt" approach: you feed the entire codebase, every dependency, and the entire task description into one massive prompt, expecting the Large Language Model (LLM) to handle the rest. This strategy, often called single-prompt bloat, is becoming increasingly expensive and brittle.

As context windows grow, the temptation to pack everything into a single inference call rises. However, the reality of LLM inference is that attention is not free. Tokenization, attention mechanisms, and the sheer volume of context lead to two primary issues:

The alternative is Multi-Agent Orchestration. This architectural pattern breaks a complex task into manageable sub-tasks, each handled by a specialized "agent" with a tightly scoped context. Tools like the gascity SDK (or similar orchestration frameworks) are designed to manage this state, routing context only to the agents that need it, thereby drastically reducing token overhead while improving code coherence.

This deep-dive explores the mechanics of why single-prompt bloat fails, how multi-agent systems solve it, and how to implement a lightweight orchestration layer to fix your AI coding workflow.

To understand why bloat is problematic, we must look at how LLMs process input. The core of a Transformer model is the Self-Attention mechanism. Mathematically, the computational cost of attention is $O(n^2)$, where $n$ is the sequence length. While recent optimizations (like Flash Attention) have mitigated the hardware constraints of $O(n^2)$, the semantic cost remains.

In a single-prompt setup, you often concatenate the following:

system_prompt: Standard instructions. global_context: The entire repository or relevant module tree. task: The user's specific request (e.g., "add a retry mechanism to the API client"). constraints: Linting rules, styling guides, etc. When the model generates the output, it must weigh every token in the sequence. If you include 50,000 tokens of context for a task that only requires 2,000 tokens of specific file data, the "signal" (the actual code to modify) is buried under "noise" (unrelated files). This leads to hallucination (inventing imports that don't exist) and inconsistency (missing global patterns because the model's attention drifted to the noisy parts).

Research has shown that LLMs perform best on information at the beginning and end of their context window. Information placed in the middle tends to be under-utilized. In bloat architectures, critical context (like the specific function to modify) is often buried in the middle of the prompt, leading to sub-par outputs that require multiple retries. Each retry burns even more tokens, creating a feedback loop of high cost and low quality.

Multi-agent orchestration shifts the paradigm from "one giant brain thinking about everything" to "a team of specialists collaborating." This approach leverages context isolation. Each agent receives only the context necessary to complete its specific sub-task.

The heart of this system is the Orchestrator. The Orchestrator does not write code; it plans. It takes the user's high-level request and decomposes it into a Directed Acyclic Graph (DAG) of tasks.

By isolating context, we avoid paying for attention across the entire repository. For example, if the "Retrieval Agent" identifies that only auth.js and utils.ts are relevant, the "Execution Agent" receives a prompt of maybe 2,000 tokens instead of 50,000. This reduction is not linear; it compounds across a team of agents.

Furthermore, because each agent's task is narrowly defined, the model can focus its attention weights on the highly relevant tokens, leading to higher code accuracy and fewer "lost in the middle" errors.

While specific SDK names vary, the gascity pattern (often associated with graph-based state management in agent workflows) emphasizes stateful routing. Let's build a conceptual implementation using a Node.js environment to demonstrate how to structure this.

We need a standard state object that travels between agents. This state holds the current context, the task history, and the specific prompt for the next step.

// types.ts
export interface AgentState {
  task: string;             // The original user request
  retrievedContext: string; // Relevant code snippets
  generatedCode: string;    // Output from execution agents
  errors: string[];         // Linting or test failures
  step: string;             // Current step in the DAG
  history: string[];        // Log of previous actions
}

export interface AgentResponse {
  nextStep: string;         // Which agent to call next
  newState: Partial<AgentState>;
}

The most critical step in saving tokens is the retrieval phase. This agent uses a semantic search engine (like a vector database) to find relevant files. It should not dump the whole repo.

// agents/retrieval.ts
import { AgentState, AgentResponse } from '../types';
import { searchCodebase } from '../services/vectorStore'; // Pseudo-code for RAG

export async function retrievalAgent(state: AgentState): Promise<AgentResponse> {
  // Use a small LLM or heuristic to determine which keywords to search
  const keywords = extractKeywords(state.task);

  // Fetch only top-K relevant snippets (e.g., K=5)
  const snippets = await searchCodebase(keywords, limit: 5);

  const context = snippets.map(s => `\n// File: ${s.path}\n${s.content}`).join('\n');

  return {
    nextStep: 'execution',
    newState: {
      retrievedContext: context,
      history: [...state.history, `Retrieved ${snippets.length} files.`]
    }
  };
}

The execution agent receives only the retrievedContext and the specific task. It does not see the rest of the repository. This forces the model to work with the provided facts, reducing hallucination.

// agents/execution.ts
import { AgentState, AgentResponse } from '../types';
import { generateCode } from '../services/llm'; // Wrapper around LLM API

export async function executionAgent(state: AgentState): Promise<AgentResponse> {
  const prompt = `\nCONTEXT:\n${state.retrievedContext}\n\nTASK:\n${state.task}\n\nOUTPUT ONLY CODE.`;

  const code = await generateCode(prompt, { temperature: 0.2 });

  return {
    nextStep: 'verification',
    newState: {
      generatedCode: code,
      history: [...state.history, 'Generated code.']
    }
  };
}

The orchestrator ties these together. It runs the agents in a loop, updating the state. If verification fails, it routes back to retrieval or execution with error feedback.

// orchestrator.ts
import { AgentState, AgentResponse } from './types';
import { retrievalAgent } from './agents/retrieval';
import { executionAgent } from './agents/execution';
import { verificationAgent } from './agents/verification';

const MAX_STEPS = 5; // Prevent infinite loops

export async function runWorkflow(task: string): Promise<AgentState> {
  let state: AgentState = {
    task: task,
    retrievedContext: '',
    generatedCode: '',
    errors: [],
    step: 'retrieval',
    history: []
  };

  for (let i = 0; i < MAX_STEPS; i++) {
    let response: AgentResponse;

    switch (state.step) {
      case 'retrieval':
        response = await retrievalAgent(state);
        break;
      case 'execution':
        response = await executionAgent(state);
        break;
      case 'verification':
        response = await verificationAgent(state); // Checks lints/tests
        break;
      default:
        throw new Error(`Unknown step: ${state.step}`);
    }

    // Update state
    state = { ...state, ...response.newState, step: response.nextStep };

    if (state.step === 'done' || state.step === 'failed') {
      break;
    }
  }

  return state;
}

Multi-agent systems are powerful but can suffer from "agent thrashing"—where agents keep disagreeing, causing infinite loops.

If the conversation history grows too long, the Orchestrator should invoke a Summarization Agent. This agent takes the last $N$ steps and compresses them into a single token-efficient summary. This summary replaces the raw history in the state, keeping the context window small.

Allow agents to use tools (e.g., read_file, run_tests). Tools are cheaper than LLM calls. If a retrieval agent can use a tool to read a specific file, it should prefer that over asking the LLM to "guess" the file content. Deterministic tools reduce token variance.

For critical tasks, the Orchestrator should before execution and present the retrievedContext to the user for approval. This prevents the system from wasting tokens on a task that is fundamentally misunderstood.

Let's estimate the token usage for a typical task: "Update the user login logic to support OAuth."

Result: Scenario B is roughly 11x more efficient in terms of input token consumption. As you scale to thousands of tasks, this difference translates to significant savings on API bills and faster inference times.

When building or adopting SDKs like gascity, keep these engineering principles in mind:

Q: Does multi-agent orchestration always use fewer tokens than a single prompt?

A: Not necessarily for trivial tasks. For a simple "fix this typo" task, a single prompt is more efficient. However, for complex, multi-file refactoring or feature implementation, multi-agent orchestration consistently reduces total token usage by avoiding context repetition and reducing retries caused by context rot.

Q: How do I prevent agents from "thrashing" (infinite loops)?

A: Implement a step limit (e.g., max 10 orchestrator cycles). Additionally, use a "progress check" where the orchestrator asks a small, cheap LLM: "Is the task closer to completion than 2 steps ago?" If the answer is no, terminate the workflow and escalate to a human.

Q: What is the role of RAG in this architecture?

A: RAG (Retrieval-Augmented Generation) is the backbone of the Retrieval Agent. It ensures that only relevant code snippets are injected into the context of the Execution Agent. Without RAG, the Execution Agent would have to "guess" the code structure, leading to hallucinations and wasted tokens on retry.

The era of "stuffing everything into the prompt" is ending. As codebases grow and AI coding assistants become more sophisticated, the ability to manage context precisely becomes a critical engineering skill. By adopting multi-agent orchestration, you not only reduce your API costs but also improve the quality of the generated code. The gascity-style approach to stateful routing and context isolation is the blueprint for the next generation of AI-assisted development workflows. Start small: break your next complex task into three agents and measure the token difference. The savings will be immediate.

── more in #ai-agents 4 stories · sorted by recency
── more on @gascity 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/stop-burning-tokens-…] indexed:0 read:8min 2026-10-09 · —