Build real-time voice applications with vLLM-Omni on SageMaker AI – Part 1 AWS published Part 1 of a tutorial series showing how to deploy the Qwen3-TTS text-to-speech model on Amazon SageMaker AI using the AWS vLLM-Omni Deep Learning Container, streaming text in and audio out over a single persistent bidirectional connection. The walkthrough uses the vLLM-Omni DLC, which packages tracked vLLM-Omni releases with routing middleware for SageMaker AI, and demonstrates the workflow through a Gradio application built from the aws-samples sagemaker-genai-hosting-examples repository. It follows an earlier post that streamed microphone audio to the Voxtral-Mini-4B Realtime speech-to-text model, adding the speech-output side of a real-time voice pipeline; Part 2 covers image and video generation with the same DLC. Artificial Intelligence https://aws.amazon.com/blogs/machine-learning/ Build real-time voice applications with vLLM-Omni on SageMaker AI – Part 1 Voice agents, interactive learning applications, accessibility tools, and customer service assistants need to respond without long silent pauses. In this tutorial, you deploy a text-to-speech TTS model on Amazon SageMaker AI https://aws.amazon.com/sagemaker/ai/ that can start playing speech before it finishes generating the full response. You use the AWS vLLM-Omni Deep Learning Container DLC https://aws.github.io/deep-learning-containers/vllm-omni/ to deploy Qwen3-TTS https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice , stream text in and audio out over one persistent bidirectional connection, and try the workflow through a Gradio application. AWS Deep Learning Containers https://github.com/aws/deep-learning-containers provide Docker images with deep learning frameworks and dependencies for training and inference on AWS. AWS provides deployment guidance for broadly adopted serving frameworks such as vLLM https://github.com/aws/deep-learning-containers/tree/master/docs/vllm and SGLang https://github.com/aws/deep-learning-containers/tree/master/docs/sglang . This post is Part 1 of a series about specialized DLCs, including vLLM-Omni https://github.com/aws/deep-learning-containers/tree/main/docs/vllm-omni , WhisperX https://github.com/aws/deep-learning-containers/tree/main/docs/whisperx , and llama.cpp https://github.com/aws/deep-learning-containers/tree/main/docs/llama-cpp . It focuses on streamed speech for real-time voice applications. Part 2 https://aws.amazon.com/blogs/machine-learning/generate-images-and-video-with-vllm-omni-on-sagemaker-ai-part-2/ applies the vLLM-Omni DLC to image and video generation. The series pairs focused use cases with deployment examples and reproducible benchmarks where they add useful evidence. The vLLM-Omni project https://github.com/vllm-project/vllm-omni extends vLLM beyond text generation to serve models that process or generate text, audio, images, and video. The AWS vLLM-Omni DLC packages tracked vLLM-Omni releases in AWS images and adds routing middleware for SageMaker AI. You use SageMaker bidirectional streaming https://docs.aws.amazon.com/sagemaker/latest/dg/realtime-endpoints-test-endpoints.html realtime-endpoints-test-endpoints-sdk to send text and receive audio chunks over the persistent connection. The earlier post, Build real-time voice applications with Amazon SageMaker AI and vLLM https://aws.amazon.com/blogs/machine-learning/build-real-time-voice-applications-with-amazon-sagemaker-ai-and-vllm/ , demonstrates the input side of a voice pipeline. It streams microphone audio to the Voxtral-Mini-4B Realtime speech-to-text STT model and returns transcription events. This post adds the output side by sending text to Qwen3-TTS and streaming generated speech back through the vLLM-Omni DLC. You will clone the code sample https://github.com/aws-samples/sagemaker-genai-hosting-examples/tree/main/03-features/bidirectional-streaming-vLLM-Omni , deploy a Qwen3-TTS model, and try streaming speech through a Gradio application. Specialized AWS DLCs for multimodal inference Specialized inference runtimes support model architectures and media pipelines that differ from general text generation. AWS DLCs package these runtimes with their framework dependencies and deployment configuration, giving you a consistent image-based path to AWS compute and managed inference services. vLLM-Omni https://github.com/vllm-project/vllm-omni extends vLLM from text-focused autoregressive generation to models that process and generate multiple modalities. Its heterogeneous pipeline abstraction coordinates multi-stage model workflows, including autoregressive and diffusion stages. It also provides streaming outputs and OpenAI-compatible APIs. The runtime is not limited to TTS. Its supported models https://github.com/vllm-project/vllm-omni/blob/main/docs/models/supported models.md include unified omni models, automatic speech recognition ASR , TTS, audio generation, image generation, and video generation. This post uses Qwen3-TTS as a focused example because streamed speech provides a direct way to demonstrate bidirectional audio output. The previous vLLM post https://aws.amazon.com/blogs/machine-learning/build-real-time-voice-applications-with-amazon-sagemaker-ai-and-vllm/ covers the input path: microphone audio enters a Voxtral model and transcription events return to the application. A voice application can pass that transcription to its conversation logic, then use this example to turn the response text into streamed speech. The orchestration between the two endpoints remains outside this walkthrough. Solution overview The complete sample lives in 03-features/bidirectional-streaming-vLLM-Omni https://github.com/aws-samples/sagemaker-genai-hosting-examples/tree/main/03-features/bidirectional-streaming-vLLM-Omni . Clone the repository to use the deployment script, shared streaming transport, and Gradio client together. SageMaker bidirectional streaming exposes a full-duplex WebSocket transported over HTTP/2. The client connects to the SageMaker Runtime endpoint on port 8443 . The SageMaker inference sidecar forwards the connection to the vLLM-Omni native WebSocket route inside the container. Figure 1 shows the request and response path. The client sends session configuration and text events. Qwen3-TTS returns audio lifecycle and chunk events over the same connection. The vLLM-Omni v1.5 DLC adds the bidirectional streaming capability label and routing middleware required by SageMaker AI. The sample uses v1/audio/speech/stream , one of the native WebSocket routes exposed by vLLM-Omni. The endpoint configuration also uses SageMaker instance pools https://docs.aws.amazon.com/sagemaker/latest/dg/realtime-endpoints-heterogeneous.html . An instance pool defines a priority-ordered list of compatible instance types for one production variant. SageMaker first tries ml.g6.xlarge , then falls back to ml.g6e.xlarge , ml.g5.xlarge , or ml.g4dn.xlarge when capacity is unavailable. It still provisions one instance, not one instance per pool. SageMaker validates quota for every configured pool entry when it creates the endpoint, so you need available quota for each included type. Your hourly cost can also change if SageMaker selects a different instance. Use --instance-types to restrict the pool to the types that meet your quota, price, and performance requirements. Prerequisites Before starting, you need: - An AWS account with credentials configured for the AWS Command Line Interface AWS CLI or an AWS SDK - Git - Python 3.12 or newer - A SageMaker AI execution role - Permission to create and invoke SageMaker AI endpoints - Endpoint quota for at least one pooled instance type: ml.g6.xlarge , ml.g6e.xlarge , ml.g5.xlarge , or ml.g4dn.xlarge - Boto3 version 1.40.0 or newer - The SageMaker Runtime HTTP/2 Python client version 0.4.0 The walkthrough uses the US East N. Virginia Region. The sample builds the DLC image URI and runtime endpoint from AWS REGION . Deploy a streaming speech endpoint 1. Clone the hosting examples repository. Clone the repository and enter the vLLM-Omni bidirectional streaming sample directory. 2. Install the required Python packages. Create a virtual environment and install the versions defined by the sample. 3. Configure the sample. Set your SageMaker AI execution role. The sample defaults to us-east-1 . Pass --region to use another supported Region.The inference AMI setting is required for the SageMaker sidecar that handles the HTTP/2 WebSocket. The model invocation path must omit the leading slash because SageMaker adds it before forwarding the request. 4. Deploy and retain the model endpoint. The sample creates a SageMaker model from the DLC, an endpoint configuration, and a real-time endpoint. The environment variable SM VLLM MODEL tells the container which model to load. The script also runs a streaming smoke test and saves the result as validation-output.wav .Wait until the endpoint reaches InService . The first deployment takes time because the instance downloads the DLC image and model artifacts. 5. Launch the Gradio application. Run the checked-in Gradio client against the endpoint.Open http://127.0.0.1:6006 in your browser. The --share option creates a Gradio public link, so leave it disabled unless you need that behavior. 6. Generate streaming speech. Enter text, choose a voice and language, and choose Generate speech . The shared client opens the bidirectional stream, configures pulse-code modulation PCM output, and sends the text.The Gradio audio component plays each 24 kHz PCM chunk as the client receives it. The status field reports the chunk count and total audio bytes. 7. Delete the sample resources. Stop the Gradio process, then delete the endpoint, endpoint configuration, and model. Clean up Confirm that the cleanup command reports deletion of the endpoint, endpoint configuration, and model. If the process stops before cleanup completes, delete the resources through the SageMaker AI console or API. A running GPU endpoint continues to incur charges. Conclusion By deploying a text-to-speech model, you used the AWS vLLM-Omni DLC on SageMaker AI to stream audio while the model was still generating it. This example shows how specialized AWS DLCs support multimodal models whose execution and streaming patterns differ from general text generation. Paired with the previous vLLM speech-to-text example https://aws.amazon.com/blogs/machine-learning/build-real-time-voice-applications-with-amazon-sagemaker-ai-and-vllm/ , the samples show complementary paths: stream microphone audio into transcription, then stream the application’s text response back as speech. Part 2 https://aws.amazon.com/blogs/machine-learning/generate-images-and-video-with-vllm-omni-on-sagemaker-ai-part-2/ which will show how to use the same DLC family for image and video generation. Try the code sample https://github.com/aws-samples/sagemaker-genai-hosting-examples/tree/main/03-features/bidirectional-streaming-vLLM-Omni , then explore the vLLM-Omni DLC documentation https://aws.github.io/deep-learning-containers/vllm-omni/ and SageMaker AI bidirectional streaming developer guide https://docs.aws.amazon.com/sagemaker/latest/dg/realtime-endpoints-test-endpoints.html realtime-endpoints-test-endpoints-sdk to apply the pattern to other supported multimodal models. Resources Acknowledgements The authors thank Zhuofu Bai, Ayush Sharma, Christian Kamwangala, and Jon Chua for their contributions to the solution and publication process.