# Synchronous vs. Asynchronous Agent Execution: Architecture Patterns for Production

> Source: <https://machinelearningmastery.com/synchronous-vs-asynchronous-agent-execution-architecture-patterns-for-production/>
> Published: 2026-10-06 11:00:55+00:00

In this article, you will learn how synchronous and asynchronous execution patterns differ architecturally, and how to choose between them when deploying LLM-based agents to production.

Topics we will cover include:

- How the synchronous “wait and see” execution pattern works, and the production scenarios where it is appropriate.
- How the asynchronous, event-driven “fire and forget” pattern decouples task submission from task completion to handle long-running agent workflows.
- A practical guide to choosing between the two patterns based on task complexity, latency tolerance, and infrastructure requirements.

## Introduction

Building local Python scripts where an **LLM-based agent** loops through several tools is increasingly easy nowadays, with endless high-level libraries and supporting frameworks on the rise. However, there is a severe “deployment gap” in the **agentic AI** space that deserves further attention.

From a production standpoint, real-world agent workflows entail intricate dependencies, multi-step reasoning loops, and API latency. Failing to design an architecture that accounts for this reality results in system timeouts, dropped user requests, and memory leaks. Two core architectural patterns that bridge these gaps and move toward standard distributed system paradigms are synchronous and asynchronous execution. This article examines these two patterns from an agent execution perspective, illustrating their rationale through two easy-to-run, lightweight notebooks whose underlying logic can be readily translated to production-ready agent-based applications.

## Synchronous Agent Execution: “Wait and See”

The synchronous agent execution pattern closely resembles the classic HTTP request-response approach. It is best suited to scenarios that demand immediate feedback, sequential dependencies between tasks, or efficient data retrieval, e.g. in standard RAG pipelines.

It works as follows: an external caller submits a prompt, then the execution thread is blocked while the agent finalizes the whole chain-of-thought process and tool usage.

While simple and effective, this pattern is limited and remarkably fragile when combined with advanced agents. Cloud infrastructure providers like AWS have a default 29-second timeout in their API Gateway. For an agent that takes 45 seconds to plan, search the web, and produce a response, a standard synchronous pipeline would drop the connection, resulting in wasted tokens and lost progress.

The following code shows a runnable example in Google Colab that emulates the synchronous agent execution pattern. It first defines a function, `mock_llm_call(prompt)`, that simulates an agent’s call to an LLM without using an actual LLM. It takes a prompt as input and simulates a small network latency, similar to invoking a real LLM. After that, simple conditional logic checks whether the prompt includes the word “search.” If it does, the function simulates a tool call for performing a web search. Otherwise, it simulates the LLM arriving at a final answer, meaning the model has processed the request and generated a direct response. In both cases, the function returns a dictionary describing an action and its associated metadata.

``` python
import time

def mock_llm_call(prompt):
    """A simulated, free LLM call (with no actual LLM used here) to demonstrate latency without API keys."""
    time.sleep(1) # Simulate network latency
    # Checking if this is the initial request requiring a search
    if "research" in prompt.lower():
        return {"action": "tool_call", "tool": "web_search", "query": "latest agent news"}
    
    return {"action": "final_answer", "text": "Found the data! Agents are scaling."}

12345678910

import time def mock_llm_call(prompt):    """A simulated, free LLM call (with no actual LLM used here) to demonstrate latency without API keys."""    time.sleep(1) # Simulate network latency    # Checking if this is the initial request requiring a search    if "research" in prompt.lower():        return {"action": "tool_call", "tool": "web_search", "query": "latest agent news"}        return {"action": "final_answer", "text": "Found the data! Agents are scaling."}
```

Meanwhile, an overarching agent function that calls the previous one is also needed. We call this function `synchronous_agent(query)`. It simulates the behavior of an agent that processes a query in several steps. At each step, the agent calls `mock_llm_call()`. If the mock LLM signals a tool call, the agent simulates the tool execution and adjusts the query for the next step. If the LLM returns a final response, the agent finalizes execution and returns that response as the final result.

``` python
def synchronous_agent(query):
    print(f"[Sync API] Blocking thread to process: '{query}'")
    max_steps = 3
    
    for step in range(max_steps):
        print(f"  -> Agent Step {step+1}: Thinking...")
        
        response = mock_llm_call(f"{query} (step {step})")
        
        if response["action"] == "final_answer":
            print(f"[Sync API] Finished! Returning payload to client.")
            return response['text']
        else:
            print(f"  -> Executing Tool: {response['tool']}")
            # Updating query to simulate injecting the tool's result, avoiding trigger words
            query = "Tool output: 'Agents require async architecture for scale.' Summarize this." 
            
    return "Agent failed to complete in time."

123456789101112131415161718

def synchronous_agent(query):    print(f"[Sync API] Blocking thread to process: '{query}'")    max_steps = 3        for step in range(max_steps):        print(f"  -> Agent Step {step+1}: Thinking...")                response = mock_llm_call(f"{query} (step {step})")                if response["action"] == "final_answer":            print(f"[Sync API] Finished! Returning payload to client.")            return response['text']        else:            print(f"  -> Executing Tool: {response['tool']}")            # Updating query to simulate injecting the tool's result, avoiding trigger words            query = "Tool output: 'Agents require async architecture for scale.' Summarize this."                 return "Agent failed to complete in time."
```

Let’s try it out:

```
# How to run the agent in a Colab notebook:
result = synchronous_agent("Research AI agent patterns")
print(f"Result: {result}")

123

# How to run the agent in a Colab notebook:result = synchronous_agent("Research AI agent patterns")print(f"Result: {result}")
```

Result:

``` php
[Sync API] Blocking thread to process: 'Research AI agent patterns'
  -> Agent Step 1: Thinking...
  -> Executing Tool: web_search
  -> Agent Step 2: Thinking...
[Sync API] Finished! Returning payload to client.
Result: Found the data! Agents are scaling.

123456

[Sync API] Blocking thread to process: 'Research AI agent patterns'  -> Agent Step 1: Thinking...  -> Executing Tool: web_search  -> Agent Step 2: Thinking...[Sync API] Finished! Returning payload to client.Result: Found the data! Agents are scaling.
```

We have just seen how a synchronous loop works in practice. The code above simulated an agent undertaking a reasoning step, calling a tool, and returning the final answer in the subsequent step.

## Asynchronous, Event-Driven Execution: “Fire and Forget”

Shifting to an asynchronous execution pattern becomes necessary when addressing complex tasks, such as refactoring codebases, multi-agent debating, or workflows that require frequent HITL (Human-In-The-Loop) approvals.

Under this pattern, task submission and task completion are fully decoupled. Once a client triggers the agent’s execution of a task, the system creates and returns a `job_id`, moving the task into a queue. Background workers then retrieve tasks from the queue and process them independently. This way, an agent can run for hours without blocking the user interface. Task states are checkpointed frequently to detect and recover from issues like node crashes, allowing the agent to resume where it left off.

The disadvantage of this pattern is the need for a more robust infrastructure, including message brokers like RabbitMQ and Redis, as well as suitable databases for modeling worker node states: PostgreSQL, MongoDB.

Let’s move on to the code example to understand the asynchronous pattern. A key building block here is Python’s `asyncio` library, whereby a “client” sends a task, receives an acknowledgment, and the task is processed in the background: ideal for time-consuming tasks that would otherwise block the user interface in a deployed application.

The `agent_worker()` function acts like a background worker node that processes tasks from a queue asynchronously, simulating long-running working steps and updating states in a database.

``` python
import asyncio
import uuid

# In-memory queue simulating Redis/Celery for free
task_queue = asyncio.Queue()
# In-memory DB simulating a persistent state checkpoint store
database = {}

async def agent_worker():
    """Background worker processing long-running agent tasks independently."""
    while True:
        task = await task_queue.get()
        task_id = task['id']
        print(f"\n[Worker] Picked up task {task_id}")
        
        # Simulating long-running multi-step agent thought process
        database[task_id] = "Running step 1 (Planning)..."
        await asyncio.sleep(1.5) 
        
        database[task_id] = "Running step 2 (Executing Tools)..."
        await asyncio.sleep(1.5)
        
        # Saving final state (checkpointing)
        database[task_id] = "Completed: Comprehensive report generated."
        print(f"[Worker] Task {task_id} finished. State saved to DB.")
        task_queue.task_done()

1234567891011121314151617181920212223242526

import asyncioimport uuid # In-memory queue simulating Redis/Celery for freetask_queue = asyncio.Queue()# In-memory DB simulating a persistent state checkpoint storedatabase = {} async def agent_worker():    """Background worker processing long-running agent tasks independently."""    while True:        task = await task_queue.get()        task_id = task['id']        print(f"\n[Worker] Picked up task {task_id}")                # Simulating long-running multi-step agent thought process        database[task_id] = "Running step 1 (Planning)..."        await asyncio.sleep(1.5)                 database[task_id] = "Running step 2 (Executing Tools)..."        await asyncio.sleep(1.5)                # Saving final state (checkpointing)        database[task_id] = "Completed: Comprehensive report generated."        print(f"[Worker] Task {task_id} finished. State saved to DB.")        task_queue.task_done()
```

Next, we have `submit_job(prompt)`. This function acts like an API that receives a task, assigns it an ID, and adds it to the queue, returning the task ID to the client immediately without waiting for the task to complete.

``` python
async def submit_job(prompt):
    """The API layer: submits job to the queue and returns immediately."""
    task_id = str(uuid.uuid4())[:8]
    await task_queue.put({"id": task_id, "prompt": prompt})
    database[task_id] = "Pending"
    return task_id

123456

async def submit_job(prompt):    """The API layer: submits job to the queue and returns immediately."""    task_id = str(uuid.uuid4())[:8]    await task_queue.put({"id": task_id, "prompt": prompt})    database[task_id] = "Pending"    return task_id
```

The last function, `main()`, is responsible for orchestrating the whole agent-based system: initiating the worker node in the background, simulating a task submission by a client, and periodically monitoring the task state until completion.

``` python
async def main():
    # 1. Starting the background worker (acting as our consumer fleet)
    worker = asyncio.create_task(agent_worker())
    
    # 2. Client submits a request (API doesn't block!)
    print("[API] Submitting heavy task...")
    job_id = await submit_job("Write a comprehensive multi-agent market report")
    print(f"[API] Success! Connection closed. Job ID returned: {job_id}\n")
    
    # 3. Client checks status periodically (Simulating frontend polling/webhooks)
    for _ in range(4):
        print(f"  [Client Polling] Status for {job_id}: {database[job_id]}")
        await asyncio.sleep(1)
        
    # Cleaning up the infinite worker for notebook safety
    worker.cancel()

# How to run natively in a Google Colab notebook (which already has a running event loop):
await main()

12345678910111213141516171819

async def main():    # 1. Starting the background worker (acting as our consumer fleet)    worker = asyncio.create_task(agent_worker())        # 2. Client submits a request (API doesn't block!)    print("[API] Submitting heavy task...")    job_id = await submit_job("Write a comprehensive multi-agent market report")    print(f"[API] Success! Connection closed. Job ID returned: {job_id}\n")        # 3. Client checks status periodically (Simulating frontend polling/webhooks)    for _ in range(4):        print(f"  [Client Polling] Status for {job_id}: {database[job_id]}")        await asyncio.sleep(1)            # Cleaning up the infinite worker for notebook safety    worker.cancel() # How to run natively in a Google Colab notebook (which already has a running event loop):await main()
```

Result:

```
[API] Submitting heavy task...
[API] Success! Connection closed. Job ID returned: 46f0c47a

  [Client Polling] Status for 46f0c47a: Pending

[Worker] Picked up task 46f0c47a
  [Client Polling] Status for 46f0c47a: Running step 1 (Planning)...
  [Client Polling] Status for 46f0c47a: Running step 2 (Executing Tools)...
[Worker] Task 46f0c47a finished. State saved to DB.
  [Client Polling] Status for 46f0c47a: Completed: Comprehensive report generated.

12345678910

[API] Submitting heavy task...[API] Success! Connection closed. Job ID returned: 46f0c47a   [Client Polling] Status for 46f0c47a: Pending [Worker] Picked up task 46f0c47a  [Client Polling] Status for 46f0c47a: Running step 1 (Planning)...  [Client Polling] Status for 46f0c47a: Running step 2 (Executing Tools)...[Worker] Task 46f0c47a finished. State saved to DB.  [Client Polling] Status for 46f0c47a: Completed: Comprehensive report generated.
```

## Conclusion: When to Use Which

As a general rule, start with a lightweight, synchronous architecture when facing fast, immediate tasks like question-answering. As your agent application grows more capable, you will typically need to transition to asynchronous execution, since the tasks it handles will take longer to process. Asynchronous systems place a higher demand on infrastructure — databases, queues, message brokers — but they are far more resilient against brittle timeouts, and that resilience is ultimately the key to scaling complex agent workflows into production.
