V-RAE: Rethinking Video Latent Spaces for Generation
Researchers propose V-RAE, a video representation autoencoder that builds compact generative latents on top of frozen vision foundation model representations, achieving 2.13 rFVD on K600 and outperforming all evaluated l…