UniServe: Serving FastH3 at Its Fastest
UniServe, a serving engine for FastH3 8-Step text-to-video-with-audio generation from the hao-ai-lab, delivers lower median end-to-end latency and 20–46% higher throughput than FastVideo, vLLM-Omni an…
UniServe, a serving engine for FastH3 8-Step text-to-video-with-audio generation from the hao-ai-lab, delivers lower median end-to-end latency and 20–46% higher throughput than FastVideo, vLLM-Omni an…
FastH3, MiniMax's video-and-audio generation model, now runs on Apple Silicon via MLX and on NVIDIA's DGX Spark desktop, with two Sparks able to generate one clip together. The release also debuts the…
FastVideo released four open-weight FastH3 Preview v0.1 checkpoints that reduce MiniMax H3 from 49 transformer calls to four, generating 768p video with 32 kHz stereo audio, while noting that fast mot…
FastVideo released FastMetal-QAD, a family of three open-source video generation models (1.3B, 5B, and 14B) that run natively on Apple Silicon via a new MLX runtime, with the 5B model generating a fiv…
Researchers introduced JetSpec, a speculative decoding method that trains a causal parallel draft head on a frozen target model to draft entire speculative trees in one pass while preserving autoregre…
JetSpec, a new speculative decoding method, trains a causal parallel draft head over fused hidden states from a frozen target model, enabling lossless verification of candidate trees in one forward pa…
FastVideo released FastWan-QAD, a family of video generation models that can produce a 5-second 480P video in 1.78 seconds on a single NVIDIA RTX 5090 using quantization-aware distillation. The models…
FastVideo has open-sourced Dreamverse, a real-time video generation workspace that enables "vibe directing" through natural-language iteration, releasing both the frontend and backend as a reference a…