Synchronous vs. Asynchronous Agent Execution: Architecture Patterns for Production A new technical guide contrasts synchronous and asynchronous execution patterns for deploying LLM-based agents to production, noting that AWS API Gateway's default 29-second timeout will drop connections for agents that take 45 seconds to plan, search the web, and respond. The guide presents two runnable Google Colab notebooks demonstrating each pattern, including a mock_llm_call function that simulates network latency with a one-second sleep, and recommends choosing between the patterns based on task complexity, latency tolerance, and infrastructure requirements. In this article, you will learn how synchronous and asynchronous execution patterns differ architecturally, and how to choose between them when deploying LLM-based agents to production. Topics we will cover include: - How the synchronous “wait and see” execution pattern works, and the production scenarios where it is appropriate. - How the asynchronous, event-driven “fire and forget” pattern decouples task submission from task completion to handle long-running agent workflows. - A practical guide to choosing between the two patterns based on task complexity, latency tolerance, and infrastructure requirements. Introduction Building local Python scripts where an LLM-based agent loops through several tools is increasingly easy nowadays, with endless high-level libraries and supporting frameworks on the rise. However, there is a severe “deployment gap” in the agentic AI space that deserves further attention. From a production standpoint, real-world agent workflows entail intricate dependencies, multi-step reasoning loops, and API latency. Failing to design an architecture that accounts for this reality results in system timeouts, dropped user requests, and memory leaks. Two core architectural patterns that bridge these gaps and move toward standard distributed system paradigms are synchronous and asynchronous execution. This article examines these two patterns from an agent execution perspective, illustrating their rationale through two easy-to-run, lightweight notebooks whose underlying logic can be readily translated to production-ready agent-based applications. Synchronous Agent Execution: “Wait and See” The synchronous agent execution pattern closely resembles the classic HTTP request-response approach. It is best suited to scenarios that demand immediate feedback, sequential dependencies between tasks, or efficient data retrieval, e.g. in standard RAG pipelines. It works as follows: an external caller submits a prompt, then the execution thread is blocked while the agent finalizes the whole chain-of-thought process and tool usage. While simple and effective, this pattern is limited and remarkably fragile when combined with advanced agents. Cloud infrastructure providers like AWS have a default 29-second timeout in their API Gateway. For an agent that takes 45 seconds to plan, search the web, and produce a response, a standard synchronous pipeline would drop the connection, resulting in wasted tokens and lost progress. The following code shows a runnable example in Google Colab that emulates the synchronous agent execution pattern. It first defines a function, mock llm call prompt , that simulates an agent’s call to an LLM without using an actual LLM. It takes a prompt as input and simulates a small network latency, similar to invoking a real LLM. After that, simple conditional logic checks whether the prompt includes the word “search.” If it does, the function simulates a tool call for performing a web search. Otherwise, it simulates the LLM arriving at a final answer, meaning the model has processed the request and generated a direct response. In both cases, the function returns a dictionary describing an action and its associated metadata. python import time def mock llm call prompt : """A simulated, free LLM call with no actual LLM used here to demonstrate latency without API keys.""" time.sleep 1 Simulate network latency Checking if this is the initial request requiring a search if "research" in prompt.lower : return {"action": "tool call", "tool": "web search", "query": "latest agent news"} return {"action": "final answer", "text": "Found the data Agents are scaling."} 12345678910 import time def mock llm call prompt : """A simulated, free LLM call with no actual LLM used here to demonstrate latency without API keys.""" time.sleep 1 Simulate network latency Checking if this is the initial request requiring a search if "research" in prompt.lower : return {"action": "tool call", "tool": "web search", "query": "latest agent news"} return {"action": "final answer", "text": "Found the data Agents are scaling."} Meanwhile, an overarching agent function that calls the previous one is also needed. We call this function synchronous agent query . It simulates the behavior of an agent that processes a query in several steps. At each step, the agent calls mock llm call . If the mock LLM signals a tool call, the agent simulates the tool execution and adjusts the query for the next step. If the LLM returns a final response, the agent finalizes execution and returns that response as the final result. python def synchronous agent query : print f" Sync API Blocking thread to process: '{query}'" max steps = 3 for step in range max steps : print f" - Agent Step {step+1}: Thinking..." response = mock llm call f"{query} step {step} " if response "action" == "final answer": print f" Sync API Finished Returning payload to client." return response 'text' else: print f" - Executing Tool: {response 'tool' }" Updating query to simulate injecting the tool's result, avoiding trigger words query = "Tool output: 'Agents require async architecture for scale.' Summarize this." return "Agent failed to complete in time." 123456789101112131415161718 def synchronous agent query : print f" Sync API Blocking thread to process: '{query}'" max steps = 3 for step in range max steps : print f" - Agent Step {step+1}: Thinking..." response = mock llm call f"{query} step {step} " if response "action" == "final answer": print f" Sync API Finished Returning payload to client." return response 'text' else: print f" - Executing Tool: {response 'tool' }" Updating query to simulate injecting the tool's result, avoiding trigger words query = "Tool output: 'Agents require async architecture for scale.' Summarize this." return "Agent failed to complete in time." Let’s try it out: How to run the agent in a Colab notebook: result = synchronous agent "Research AI agent patterns" print f"Result: {result}" 123 How to run the agent in a Colab notebook:result = synchronous agent "Research AI agent patterns" print f"Result: {result}" Result: php Sync API Blocking thread to process: 'Research AI agent patterns' - Agent Step 1: Thinking... - Executing Tool: web search - Agent Step 2: Thinking... Sync API Finished Returning payload to client. Result: Found the data Agents are scaling. 123456 Sync API Blocking thread to process: 'Research AI agent patterns' - Agent Step 1: Thinking... - Executing Tool: web search - Agent Step 2: Thinking... Sync API Finished Returning payload to client.Result: Found the data Agents are scaling. We have just seen how a synchronous loop works in practice. The code above simulated an agent undertaking a reasoning step, calling a tool, and returning the final answer in the subsequent step. Asynchronous, Event-Driven Execution: “Fire and Forget” Shifting to an asynchronous execution pattern becomes necessary when addressing complex tasks, such as refactoring codebases, multi-agent debating, or workflows that require frequent HITL Human-In-The-Loop approvals. Under this pattern, task submission and task completion are fully decoupled. Once a client triggers the agent’s execution of a task, the system creates and returns a job id , moving the task into a queue. Background workers then retrieve tasks from the queue and process them independently. This way, an agent can run for hours without blocking the user interface. Task states are checkpointed frequently to detect and recover from issues like node crashes, allowing the agent to resume where it left off. The disadvantage of this pattern is the need for a more robust infrastructure, including message brokers like RabbitMQ and Redis, as well as suitable databases for modeling worker node states: PostgreSQL, MongoDB. Let’s move on to the code example to understand the asynchronous pattern. A key building block here is Python’s asyncio library, whereby a “client” sends a task, receives an acknowledgment, and the task is processed in the background: ideal for time-consuming tasks that would otherwise block the user interface in a deployed application. The agent worker function acts like a background worker node that processes tasks from a queue asynchronously, simulating long-running working steps and updating states in a database. python import asyncio import uuid In-memory queue simulating Redis/Celery for free task queue = asyncio.Queue In-memory DB simulating a persistent state checkpoint store database = {} async def agent worker : """Background worker processing long-running agent tasks independently.""" while True: task = await task queue.get task id = task 'id' print f"\n Worker Picked up task {task id}" Simulating long-running multi-step agent thought process database task id = "Running step 1 Planning ..." await asyncio.sleep 1.5 database task id = "Running step 2 Executing Tools ..." await asyncio.sleep 1.5 Saving final state checkpointing database task id = "Completed: Comprehensive report generated." print f" Worker Task {task id} finished. State saved to DB." task queue.task done 1234567891011121314151617181920212223242526 import asyncioimport uuid In-memory queue simulating Redis/Celery for freetask queue = asyncio.Queue In-memory DB simulating a persistent state checkpoint storedatabase = {} async def agent worker : """Background worker processing long-running agent tasks independently.""" while True: task = await task queue.get task id = task 'id' print f"\n Worker Picked up task {task id}" Simulating long-running multi-step agent thought process database task id = "Running step 1 Planning ..." await asyncio.sleep 1.5 database task id = "Running step 2 Executing Tools ..." await asyncio.sleep 1.5 Saving final state checkpointing database task id = "Completed: Comprehensive report generated." print f" Worker Task {task id} finished. State saved to DB." task queue.task done Next, we have submit job prompt . This function acts like an API that receives a task, assigns it an ID, and adds it to the queue, returning the task ID to the client immediately without waiting for the task to complete. python async def submit job prompt : """The API layer: submits job to the queue and returns immediately.""" task id = str uuid.uuid4 :8 await task queue.put {"id": task id, "prompt": prompt} database task id = "Pending" return task id 123456 async def submit job prompt : """The API layer: submits job to the queue and returns immediately.""" task id = str uuid.uuid4 :8 await task queue.put {"id": task id, "prompt": prompt} database task id = "Pending" return task id The last function, main , is responsible for orchestrating the whole agent-based system: initiating the worker node in the background, simulating a task submission by a client, and periodically monitoring the task state until completion. python async def main : 1. Starting the background worker acting as our consumer fleet worker = asyncio.create task agent worker 2. Client submits a request API doesn't block print " API Submitting heavy task..." job id = await submit job "Write a comprehensive multi-agent market report" print f" API Success Connection closed. Job ID returned: {job id}\n" 3. Client checks status periodically Simulating frontend polling/webhooks for in range 4 : print f" Client Polling Status for {job id}: {database job id }" await asyncio.sleep 1 Cleaning up the infinite worker for notebook safety worker.cancel How to run natively in a Google Colab notebook which already has a running event loop : await main 12345678910111213141516171819 async def main : 1. Starting the background worker acting as our consumer fleet worker = asyncio.create task agent worker 2. Client submits a request API doesn't block print " API Submitting heavy task..." job id = await submit job "Write a comprehensive multi-agent market report" print f" API Success Connection closed. Job ID returned: {job id}\n" 3. Client checks status periodically Simulating frontend polling/webhooks for in range 4 : print f" Client Polling Status for {job id}: {database job id }" await asyncio.sleep 1 Cleaning up the infinite worker for notebook safety worker.cancel How to run natively in a Google Colab notebook which already has a running event loop :await main Result: API Submitting heavy task... API Success Connection closed. Job ID returned: 46f0c47a Client Polling Status for 46f0c47a: Pending Worker Picked up task 46f0c47a Client Polling Status for 46f0c47a: Running step 1 Planning ... Client Polling Status for 46f0c47a: Running step 2 Executing Tools ... Worker Task 46f0c47a finished. State saved to DB. Client Polling Status for 46f0c47a: Completed: Comprehensive report generated. 12345678910 API Submitting heavy task... API Success Connection closed. Job ID returned: 46f0c47a Client Polling Status for 46f0c47a: Pending Worker Picked up task 46f0c47a Client Polling Status for 46f0c47a: Running step 1 Planning ... Client Polling Status for 46f0c47a: Running step 2 Executing Tools ... Worker Task 46f0c47a finished. State saved to DB. Client Polling Status for 46f0c47a: Completed: Comprehensive report generated. Conclusion: When to Use Which As a general rule, start with a lightweight, synchronous architecture when facing fast, immediate tasks like question-answering. As your agent application grows more capable, you will typically need to transition to asynchronous execution, since the tasks it handles will take longer to process. Asynchronous systems place a higher demand on infrastructure — databases, queues, message brokers — but they are far more resilient against brittle timeouts, and that resilience is ultimately the key to scaling complex agent workflows into production.