{"slug": "megaslide-dit-memory-centric-adaptation-and-deformable-local-attention-for-video", "title": "MegaSlide-DiT: Memory-Centric Adaptation and Deformable Local Attention for Efficient Video Diffusion", "summary": "Researchers introduce MegaSlide-DiT, a prototype that adapts a pre-trained 105-billion-parameter Diffusion Transformer (DiT) for high-resolution video generation on a single Nvidia H200 GPU with 1.5 TB of host RAM, by keeping persistent model state in host memory and streaming transient shards to the GPU. The system replaces quadratic global attention with 3D Deformable Slide Attention (3D-DSA), reducing memory and computational complexity to linear in sequence length, offering a pragmatic path for full-parameter adaptation of massive video diffusion models without large GPU clusters.", "body_md": "arXiv:2607.22696v1 Announce Type: new\nAbstract: High-resolution video diffusion models built on Diffusion Transformers (DiTs) deliver strong fidelity but quickly exhaust the memory budget of a single workstation. A 100 billion-plus parameter DiT easily requires over a terabyte of persistent state, while naive spatiotemporal self-attention grows quadratically in sequence length. These two walls -- parameter memory and activation memory -- prevent researchers from adapting massive generative models without large GPU clusters. We revisit this problem from a systems perspective and introduce MegaSlide-DiT, a prototype that demonstrates how a pre-trained 105B DiT can be adapted on a single H200 GPU with 1.5 TB of host RAM. Our key insight is that the GPU need not own the model state: all persistent weights, master weights and optimizer moments remain in host memory, while only transient shards are streamed to the GPU on demand. Simultaneously, we replace quadratic global attention with 3D Deformable Slide Attention (3D-DSA), a motion-adaptive local attention operator that reduces both memory and computational complexity to linear in the sequence length. We report detailed memory accounting, execution traces and evaluation results to substantiate our design. MegaSlide-DiT does not claim to train a 105B model from scratch on a single GPU, nor does it magically solve bandwidth limits; rather, it offers a pragmatic path for full-parameter adaptation of massive video diffusion models on high-end workstations.", "url": "https://wpnews.pro/news/megaslide-dit-memory-centric-adaptation-and-deformable-local-attention-for-video", "canonical_source": "https://arxiv.org/abs/2607.22696", "published_at": "2026-07-28 04:00:00+00:00", "updated_at": "2026-07-28 04:07:52.202877+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-infrastructure", "computer-vision"], "entities": ["MegaSlide-DiT", "Diffusion Transformer (DiT)", "Nvidia H200", "3D Deformable Slide Attention (3D-DSA)"], "alternates": {"html": "https://wpnews.pro/news/megaslide-dit-memory-centric-adaptation-and-deformable-local-attention-for-video", "markdown": "https://wpnews.pro/news/megaslide-dit-memory-centric-adaptation-and-deformable-local-attention-for-video.md", "text": "https://wpnews.pro/news/megaslide-dit-memory-centric-adaptation-and-deformable-local-attention-for-video.txt", "jsonld": "https://wpnews.pro/news/megaslide-dit-memory-centric-adaptation-and-deformable-local-attention-for-video.jsonld"}}