cd /news/artificial-intelligence/sana-video-2-0 · home topics artificial-intelligence article
[ARTICLE · art-71570] src=nvlabs.github.io ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

SANA-Video 2.0

NVIDIA Research's Efficient AI Team and Singapore Lab introduced SANA-Video 2.0, a hybrid video diffusion transformer at 5B and 14B scales that generates 720p video on a single H100 GPU. The model achieves a VBench score of 84.30 in 13.06 seconds for 720p/5s, is 3.2× faster than a full-softmax baseline at 720p/60s, and 120× faster than Wan 2.2 14B on one H100, using hybrid linear-softmax attention and block attention residuals to match full-softmax quality with linear attention efficiency.

read2 min views1 publishedJul 24, 2026

NVIDIA Research · Efficient AI Team & Singapore Lab

Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

84.30 VBench Total

13.06s 720p/5s · one H100

3.2× faster than softmax at 60 s

120× faster than Wan 2.2 14B

One-H100 latency #

720p / 5s · one H100 · 40 steps

120×

1556

788

130

69.3

13.06

log scaleseconds ↓

Paper overview

Abstract #

Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video DiTs in quality while retaining the favorable long-sequence scaling of linear attention. To avoid quadratic attention throughout, Hybrid Linear-Softmax Attention combines gated linear attention for O(N)-dominated mixing with periodic gated-softmax anchors at a 3:1 ratio, restoring the full-rank token interactions that pure linear attention lacks. To propagate these refreshed representations across depth, Block Attention Residuals (AttnRes) route completed block summaries into later linear layers, enabling anchor-feature reuse and boosting deep-layer effective rank by ~12%. Through from-scratch training, SANA-Video 2.0 learns the complete hybrid directly rather than linearizing pretrained models, with reduced-resolution proxy studies establishing 25% softmax as the optimal quality-efficiency trade-off. With 40-step sampling, SANA-Video 2.0 achieves a VBench score of 84.30 in 13.2s at 480p on a single H100, remaining competitive with far larger softmax video DiTs at a fraction of the latency. Its compiled DiT forward pass is 3.2× faster than a matched full-softmax baseline at 720p/60s, a gap that expands with video duration. Furthermore, full-stack Sol-Engine optimization (kernel fusion, caching, and sparse attention) accelerates this hardware-friendly backbone by a further 3.58×, bringing the 5B pipeline to 13.06s at 720p/5s and making it 120× faster than Wan 2.2-A14B on one H100. Overall, our hybrid design recovers softmax-level expressiveness at substantially reduced cost, unlocking scalable long, high resolution video generation.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @nvidia research 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/sana-video-2-0] indexed:0 read:2min 2026-07-24 ·