Zero-Shot Video Restoration and Enhancement with Text-to-Image Latent Diffusion Models and Multi-Modal References Researchers propose a zero-shot video restoration and enhancement framework using a text-to-image latent diffusion model and multi-modal references, reducing inference time to nearly one-third of the original while improving temporal consistency. The method, detailed in an arXiv paper (2608.26476v1), introduces dual prompt tuning inversion and sampling, texture-aware video token merging, and referenced self-attention to support image reference, outperforming existing approaches in restoring temporally consistent videos. arXiv:2608.26476v1 Announce Type: new Abstract: Zero-shot image restoration methods with text-to-image latent diffusion models have achieved great success in universal image restoration tasks without training. However, applying them to video restoration will result in severe temporal flickering. In this paper, we propose a novel framework for zero-shot video restoration and enhancement which uses a text-to-image latent diffusion model and multi-modal references. Through the proposed dual prompt tuning inversion and sampling, the inference time can be reduced to nearly 1/3 of the original. The performance and temporal consistency can be also significantly stregthened. By using the proposed texture-aware video token merging, the temporal correlation between frames can be further utilized to improve the temporal consistency. We futher propose the referenced self-attention and referenced token merging to support image reference. Experimental results demonstrate the superiority of the proposed method in restoring and enhancing temporally consistent videos.