Don't Block Your GPU: Architecting a Distributed AI Audio Backend with FastAPI, Celery, and Redis A developer detailed an architecture for scaling AI audio transcription backends by decoupling FastAPI's web layer from GPU inference using Celery and Redis. The writeup shows that running synchronous Faster-Whisper transcription inside a FastAPI endpoint blocks the event loop and causes CUDA out-of-memory errors under concurrent load, and proposes routing GPU-bound work to a dedicated gpu_tasks queue and I/O-bound jobs like PDF generation to a cpu_tasks queue. The project, SoundPulse, uses separate hardware-optimized worker pools to avoid VRAM fragmentation and keep the GPU from idling behind disk-bound tasks. You’ve just built a brilliant AI audio transcription system using PyTorch and Faster-Whisper. It works flawlessly in your local environment. To show it to the world, you wrap it in a FastAPI endpoint, deploy it, and send a test request. It works perfectly. Then, three users upload audio files at the exact same time. Suddenly, your server freezes, new requests time out, and your terminal spits out the dreaded RuntimeError: CUDA out of memory What just happened? Welcome to the Event Loop Trap . Here is the exact anti-pattern that causes this nightmare: python @app.post "/transcribe" async def transcribe audio file: UploadFile : THE TRAP: This synchronous, compute-bound task blocks the entire event loop result = whisper model.transcribe file.file return {"text": result} FastAPI is blazing fast because it relies on an asynchronous event loop to handle concurrent requests. However, machine learning inference—whether it is Whisper, LLaMA, or Qwen—is heavily synchronous and compute-bound. When you place a 30-second audio transcription task directly inside an async def or even a standard def endpoint, you are essentially hijacking the server's main thread. While your GPU is crunching matrices, FastAPI is completely paralyzed. It cannot accept new connections, it cannot return health checks, and it forces incoming requests to wait until they eventually time out. It is the architectural equivalent of having a single cashier at a busy supermarket who stops taking customers to go manually restock the shelves. To build a production-ready AI backend, we need to completely decouple the web layer from the hardware layer. We need to stop making our users wait, and start using asynchronous task queues The golden rule of building scalable web APIs is simple: Never perform heavy computation within the request-response cycle. Instead of forcing FastAPI to wait for the GPU to finish its job, we need to transform our API into a lightweight Gateway. Its only responsibility should be to receive the audio file, hand it over to a background worker, and immediately return a task id to the client. The client can then use this ID to poll for the status via a simple REST endpoint. To achieve this structural decoupling, we introduce two critical components: Redis The Message Broker : Acts as the reliable middleman. When FastAPI receives a request, it publishes a message to Redis saying, "Hey, there is a new audio file waiting to be transcribed." Celery The Task Queue : A fleet of background workers constantly listening to Redis. When a message arrives, a worker picks it up, processes the audio on the hardware, and saves the final result. In a real-world system like SoundPulse https://github.com/KadirCanCelik/SoundPulse , transcription isn't the only task. After the AI model extracts the text, we might need to generate PDF reports, parse metadata, or perform heavy database updates. If we throw all these tasks into a single queue, we create a new, hidden bottleneck. A lightweight PDF generation task might get stuck in line behind a massive 1-hour audio transcription. Even worse, your expensive GPU might sit completely idle while a worker is busy formatting a PDF document. The architectural solution is establishing Dedicated Queues: gpu tasks Queue: Strictly reserved for PyTorch and Faster-Whisper. These workers demand heavy VRAM and parallel compute. cpu tasks Queue: Reserved for I/O-bound operations like PDF report generation and database transactions. By routing tasks to their respective hardware-optimized workers, we ensure maximum resource utilization. The GPU never waits for the disk, and the disk never waits for the GPU. Setting up separate queues is only half the battle. If you start your GPU-bound Celery worker with the default settings, you will immediately run into a wall: VRAM Fragmentation and CUDA Out of Memory OOM errors. By default, Celery uses a prefork execution pool, meaning it forks multiple child processes based on your CPU cores. If you have an 8-core machine, Celery will try to spawn 8 concurrent worker processes. Now, imagine 8 different processes simultaneously trying to load a massive PyTorch model like Faster-Whisper into your hardware's limited VRAM. Your GPU will crash instantly. To prevent the framework from aggressively multiplexing the GPU, we must force the worker to process exactly one task at a time in the main thread. We achieve this by starting the Celery worker with the --pool=solo flag: GPU worker command: celery -A src.worker.worker worker -Q gpu queue --pool=solo --loglevel=info --max-tasks-per-child=20 -E CPU worker command: celery -A src.worker.worker worker -Q cpu queue --concurrency=4 --loglevel=info --max-tasks-per-child=50 -E --pool=solo GPU only : Forces the worker to process exactly one task at a time in the main thread. It prevents the framework from aggressively multiplexing the GPU, keeping VRAM usage predictable. --max-tasks-per-child: The absolute lifesaver. It forces the worker process to gracefully die after a set number of tasks 20 for GPU tasks, 50 for CPU tasks and spawns a fresh process in its place to prevent gradual memory leaks from accumulating in long-running workers. With our queues configured and our hardware protected, let's look at how our FastAPI endpoint becomes a high-speed traffic director orchestrating a multi-stage pipeline. Instead of just triggering one task, we use Celery's chain to create a workflow. The audio is first transcribed on the heavily restricted gpu queue . Once finished, the result is automatically piped to the cpu queue for analysis—without any intervention from FastAPI. To allow the frontend to track this multi-stage process seamlessly, we generate a Composite Job ID task a id::task b id . Here is the exact production code: python @router.post "/analyze" async def analyze audio file: UploadFile = File ... , num speakers: int = Form default=2, description="Expected number of speakers" : valid extensions = ".wav", ".mp3", ".m4a", ".flac" if not file.filename.lower .endswith valid extensions : raise HTTPException status code=400, detail="Unsupported audio format." try: safe filename = file.filename.replace " ", " " file path = os.path.join UPLOAD DIR, safe filename with open file path, "wb" as buffer: shutil.copyfileobj file.file, buffer task chain = chain celery client.signature 'src.worker.tasks.transcribe audio task', args= file path, num speakers, file.filename .set queue='gpu queue' , celery client.signature 'src.worker.tasks.analyze and report task' .set queue='cpu queue' task result = task chain.apply async task a id = task result.parent.id if task result.parent else "none" task b id = task result.id composite job id = f"{task a id}::{task b id}" return { "message": "Analysis pipeline chained and queued successfully.", "job id": composite job id, "status": "PENDING" } except Exception as e: logger.error f"Analysis routing failed: {e}", exc info=True raise HTTPException status code=500, detail=str e @router.get "/status/{job id}" async def get status job id: str : """ Endpoint for the UI to poll the real-time processing status. Uses Composite ID to track both GPU and CPU tasks flawlessly. """ if "::" in job id: task a id, task b id = job id.split "::" else: task a id = task b id = job id node b = celery client.AsyncResult task b id active node = node b if node b.state == 'PENDING' and task a id = "none": node a = celery client.AsyncResult task a id if node a.state in 'PENDING', 'STARTED', 'PROCESSING', 'FAILED' : active node = node a elif node a.state == 'SUCCESS': return { "job id": job id, "status": "PROCESSING", "result": None, "meta": {"step": "GPU processing completed. Initializing Semantic Analysis...", "progress": 45} } response = { "job id": job id, "status": active node.state, "result": None, "meta": None } if active node.state == 'SUCCESS': response "result" = active node.result elif active node.state == 'PROCESSING': response "meta" = active node.info elif active node.state == 'FAILED': response "result" = str active node.info return response The era of wrapping a local Python script in a FastAPI endpoint and calling it a "production AI system" is over. True engineering in the AI space isn't just about invoking LLM APIs or loading model weights; it is about rigorous system design. It is about understanding the strict limits of your hardware, decoupling your web layer from heavy computation, and guaranteeing that your event loop never blocks. By introducing a message broker like Redis and a task queue like Celery, we transformed a fragile, monolithic API into a resilient, distributed data pipeline. The next time you architect an AI backend, ask yourself:"Is my architecture efficiently serving the ML model, or is the model bottlenecking my entire architecture?" Don't block your GPU. Decouple your queues, chain your workflows, protect your hardware, and engineer systems that actually scale. This article covers the core architectural logic, but the complete implementation of SoundPulse https://github.com/KadirCanCelik/SoundPulse is entirely open-source.