Sequential Pipelines Are Killing Your Agent Throughput: Concurrent Execution Patterns That Cut Latency by 3x A developer outlined concurrent execution patterns for multi-agent LLM pipelines, arguing that sequential orchestration leaves agents idle and inflates latency. The writeup shows that independent agent calls run with Promise.all complete in the time of the slowest task rather than the sum, and presents Promise.allSettled and directed-acyclic-graph scheduling as primitives for partial results and mixed dependencies. A user abandoned a research report at second 41. The pipeline was still running. Five agents, each waiting for the previous one to finish, even though three of them had no actual dependency on each other. By the time the output landed, the tab was closed. Sequential orchestration looks reasonable on a whiteboard. Planner decomposes the query. Retriever fetches sources. Analyzer produces findings. Verifier checks claims. Synthesizer merges everything. Five stages, each logically dependent on the last. In production, it is a serial queue of LLM calls. The retriever cannot start until the planner finishes. The analyzer cannot start until every retrieval completes. The verifier waits on the analyzer. The synthesizer waits on everything. Every agent is idle while the previous one finishes work that has no actual dependency on it. If market analysis takes 12 seconds and risk assessment takes 10 seconds, sequential execution takes 22 seconds. Parallel execution completes in 12 seconds the maximum of the two, not the sum . For workflows with independent tasks, parallel execution cuts latency proportionally to the number of concurrent agents. Most agent pipelines have fewer dependencies than their sequential structure implies. The problem is not knowing which tasks can run in parallel without breaking correctness. Start by drawing the actual dependency graph: This looks like a strict chain. But zoom in on the retriever. If the planner produces four sub-questions, the retriever can fetch sources for all four simultaneously. If the analyzer produces findings for each sub-question, those analyses can run in parallel. If the verifier checks claims independently, those checks can run concurrently. The real dependency graph has three parallelization points: A sequential pipeline treats these as single-threaded loops. A concurrent pipeline treats them as parallel map operations. You need three primitives to orchestrate concurrent agent execution: When multiple agent calls have no dependencies on each other, launch them all and wait for the slowest one: js // Sequential: 40 seconds total const marketAnalysis = await agent.analyze 'market trends' ; const riskAssessment = await agent.analyze 'risk factors' ; const competitorScan = await agent.analyze 'competitor landscape' ; const regulatoryCheck = await agent.analyze 'regulatory environment' ; // Concurrent: 12 seconds slowest agent const marketAnalysis, riskAssessment, competitorScan, regulatoryCheck = await Promise.all agent.analyze 'market trends' , agent.analyze 'risk factors' , agent.analyze 'competitor landscape' , agent.analyze 'regulatory environment' ; This works when tasks share no state and produce independent outputs. If one agent call fails, the entire batch fails. That is the correct behavior for truly independent work. When you need results from successful agents even if some fail: js const results = await Promise.allSettled agent.analyze 'market trends' , agent.analyze 'risk factors' , agent.analyze 'competitor landscape' , agent.analyze 'regulatory environment' ; const successful = results .filter r = r.status === 'fulfilled' .map r = r.value ; const failed = results .filter r = r.status === 'rejected' .map r = { reason: r.reason } ; // Proceed with partial results or retry only failed tasks This pattern is critical when agent calls hit rate limits, timeouts, or transient API failures. You get partial results immediately and can retry only the failed subset. When tasks have mixed dependencies some parallel, some sequential , model the workflow as a directed acyclic graph: js const dag = { planner: { deps: , fn: = agent.plan query }, retriever: { deps: 'planner' , fn: plan = agent.retrieve plan.subQuestions }, analyzer: { deps: 'retriever' , fn: sources = agent.analyze sources }, verifier: { deps: 'analyzer' , fn: findings = agent.verify findings }, synthesizer: { deps: 'verifier' , fn: validated = agent.synthesize validated } }; async function executeDag dag { const results = {}; const completed = new Set ; while completed.size < Object.keys dag .length { const ready = Object.entries dag .filter name, node = completed.has name && node.deps.every dep = completed.has dep ; const batch = await Promise.all ready.map async name, node = { const depResults = node.deps.map dep = results dep ; results name = await node.fn ...depResults ; completed.add name ; } ; } return results; } This executor runs all tasks with satisfied dependencies in parallel. The planner runs first. Once it completes, the retriever runs. Once the retriever completes, all analyzer calls run concurrently. The DAG shape determines the parallelism automatically. Sequential pipelines have simple observability. One span per agent call, nested in a linear trace. Concurrent pipelines require different instrumentation. Each parallel batch becomes a single parent span with multiple child spans: research report 45s → 15s ├─ planner 3s └─ retrieval batch 12s ├─ retrieve market 12s ├─ retrieve risk 8s ├─ retrieve competitor 10s └─ retrieve regulatory 7s └─ analysis batch 8s ├─ analyze market 8s ├─ analyze risk 6s ├─ analyze competitor 7s └─ analyze regulatory 5s └─ verification batch 4s ├─ verify claim 1 4s ├─ verify claim 2 3s └─ verify claim 3 2s └─ synthesizer 2s The parent span duration is now the maximum of its children, not the sum. This makes waterfall views accurate. Track these for concurrent execution: When five agents run concurrently and one fails, you need to know which one without replaying the entire batch. Tag each agent call with: This lets you retry a single failed agent call without re-running the entire batch. Concurrent execution introduces failure modes that do not exist in sequential pipelines. | Failure Mode | Cause | Mitigation | |---|---|---| | Rate limit cascade | Five agents hit the same API simultaneously, all get 429s | Semaphore to limit concurrent calls per provider | | Memory spike | All agents load large context simultaneously | Stream responses or limit batch size | | Partial state corruption | One agent writes shared state while another reads it | Immutable inputs, copy-on-write outputs | | Deadlock | Agent A waits for B, B waits for A cycle in DAG | Topological sort before execution | | Straggler amplification | One slow agent blocks an entire batch | Timeout per agent, proceed with partial results | | Observability explosion | 5x more spans overwhelm your tracing backend | Sample traces or aggregate batch spans | The most common failure is rate limit cascade. If your LLM provider allows 10 requests per second and you launch 20 concurrent agent calls, half will fail immediately. Use a semaphore to limit concurrency: js const semaphore = new Semaphore 10 ; // Max 10 concurrent calls async function rateLimitedCall fn { await semaphore.acquire ; try { return await fn ; } finally { semaphore.release ; } } const results = await Promise.all tasks.map task = rateLimitedCall = agent.call task ; This queues excess calls instead of failing them immediately. Concurrent execution is not always faster. Use sequential orchestration when: The decision is not about performance. It is about whether your workflow has actual dependencies or just appears to have them because you wrote it as a loop. Use concurrent execution when: Avoid concurrent execution when: The 3x improvement comes from running independent work in parallel instead of in sequence. If your pipeline has no independent work, concurrent execution will not help. Draw the dependency graph first. If it is a straight line, stay sequential. If it has branches, parallelize them.