{"slug": "sana-video-2-0", "title": "SANA-Video 2.0", "summary": "NVIDIA Research's Efficient AI Team and Singapore Lab introduced SANA-Video 2.0, a hybrid video diffusion transformer at 5B and 14B scales that generates 720p video on a single H100 GPU. The model achieves a VBench score of 84.30 in 13.06 seconds for 720p/5s, is 3.2× faster than a full-softmax baseline at 720p/60s, and 120× faster than Wan 2.2 14B on one H100, using hybrid linear-softmax attention and block attention residuals to match full-softmax quality with linear attention efficiency.", "body_md": "NVIDIA Research · Efficient AI Team & Singapore Lab\n\n# SANA-Video 2.0\n\nHybrid Linear Attention with Attention Residuals for Efficient Video Generation\n\n**84.30** VBench Total\n\n**13.06s** 720p/5s · one H100\n\n**3.2×** faster than softmax at 60 s\n\n**120×** faster than Wan 2.2 14B\n\n## One-H100 latency\n\n720p / 5s · one H100 · 40 steps\n\n**120×**\n\n**1556**\n\n**788**\n\n**130**\n\n**69.3**\n\n**13.06**\n\nlog scaleseconds ↓\n\nPaper overview\n\n## Abstract\n\nHybrid Linear Attention with Attention Residuals for Efficient Video Generation\n\nWe introduce **SANA-Video 2.0**, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video DiTs in quality while retaining the favorable long-sequence scaling of linear attention. To avoid quadratic attention throughout, **Hybrid Linear-Softmax Attention** combines gated linear attention for O(N)-dominated mixing with periodic gated-softmax anchors at a 3:1 ratio, restoring the full-rank token interactions that pure linear attention lacks. To propagate these refreshed representations across depth, **Block Attention Residuals (AttnRes)** route completed block summaries into later linear layers, enabling anchor-feature reuse and boosting deep-layer effective rank by ~12%. Through from-scratch training, SANA-Video 2.0 learns the complete hybrid directly rather than linearizing pretrained models, with reduced-resolution proxy studies establishing 25% softmax as the optimal quality-efficiency trade-off. With 40-step sampling, SANA-Video 2.0 achieves a VBench score of 84.30 in 13.2s at 480p on a single H100, remaining competitive with far larger softmax video DiTs at a fraction of the latency. Its compiled DiT forward pass is **3.2× faster than a matched full-softmax baseline** at 720p/60s, a gap that expands with video duration. Furthermore, full-stack Sol-Engine optimization (kernel fusion, caching, and sparse attention) accelerates this hardware-friendly backbone by a further **3.58×**, bringing the 5B pipeline to **13.06s at 720p/5s** and making it **120× faster than Wan 2.2-A14B** on one H100. Overall, our hybrid design recovers softmax-level expressiveness at substantially reduced cost, unlocking scalable long, high resolution video generation.", "url": "https://wpnews.pro/news/sana-video-2-0", "canonical_source": "https://nvlabs.github.io/Sana/Video2/", "published_at": "2026-07-24 06:34:56+00:00", "updated_at": "2026-07-24 06:52:37.869145+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "computer-vision", "ai-research", "ai-infrastructure"], "entities": ["NVIDIA Research", "Efficient AI Team", "Singapore Lab", "SANA-Video 2.0", "H100", "Wan 2.2 14B", "Sol-Engine"], "alternates": {"html": "https://wpnews.pro/news/sana-video-2-0", "markdown": "https://wpnews.pro/news/sana-video-2-0.md", "text": "https://wpnews.pro/news/sana-video-2-0.txt", "jsonld": "https://wpnews.pro/news/sana-video-2-0.jsonld"}}