Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation Video DeltaNet introduces a video-native hybrid attention mechanism for livestream video generation, targeting the computational bottleneck caused by repeated processing of long spatiotemporal token sequences during denoising in video diffusion models. The approach draws on linear attention, which has been widely adopted in recent large language models, though the source notes that directly applying linear attention to video models presents challenges. Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck. Linear attention offers an appealing alternative and has been widely adopted in recent large language models, but directly applying it to video models of