UniServe: Serving FastH3 at Its Fastest
UniServe, a serving engine for FastH3 8-Step text-to-video-with-audio generation from the hao-ai-lab, delivers lower median end-to-end latency and 20–46% higher throughput than FastVideo, vLLM-Omni an…
UniServe, a serving engine for FastH3 8-Step text-to-video-with-audio generation from the hao-ai-lab, delivers lower median end-to-end latency and 20–46% higher throughput than FastVideo, vLLM-Omni an…
AWS published Part 1 of a tutorial series showing how to deploy the Qwen3-TTS text-to-speech model on Amazon SageMaker AI using the AWS vLLM-Omni Deep Learning Container, streaming text in and audio o…
AWS published Part 2 of its vLLM-Omni on SageMaker AI series, showing how to deploy two endpoints from the same AWS vLLM-Omni Deep Learning Container to generate images and video. A SageMaker AI real-…
OpenBMB released VoxCPM2, a 2-billion-parameter, Apache-2.0 licensed text-to-speech model that clones voices from a few seconds of audio, generates new voices from text descriptions, and outputs 48 kH…
Alibaba's Qwen released Qwen-Image-2.1, a 7B-parameter image generation and editing model that combines a 32-layer single-stream DiT with a Qwen3-VL 8B text encoder and a 64-channel RGBA VAE, but ship…
MiniMax promoted an H3 integration index on August 25, cataloging community tools for local deployment on 24GB GPUs, ComfyUI nodes, and multi-GPU serving via SGLang and vLLM-Omni, three weeks after re…
Nari Labs' Qwen3-TTS 1.7B CustomVoice implementation achieves 10 requests per second and sub-50 ms p95 time-to-first-audio on a single NVIDIA H100 SXM, maintaining real-time playback and costing about…
Nika is a new open-source workflow language for AI that turns repeatable AI tasks into portable, auditable files. The tool, built as a single Rust binary, allows users to define workflows with four ve…