{"slug": "video-models-as-native-4d-renderers-world-grounded-conditioning-from-animated", "title": "Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh", "summary": "Researchers propose DAR, a reference-guided renderer that extends Wan2.2 camera control to a joint camera-plus-geometry interface, enabling pretrained video diffusion models to act as 4D renderers from animated meshes. On the 68-case DAR-4D benchmark, LoRA DAR achieves PSNR 23.22, SSIM 0.895, and LPIPS 0.134, improving over off-the-shelf Wan2.2-Depth by 1.54 dB PSNR, while a full fine-tune reaches PSNR 25.36 and SSIM 0.917. The work demonstrates that tracking and world position, rather than depth, are effective conditions for 4D rendering.", "body_md": "arXiv:2608.00094v1 Announce Type: new\nAbstract: Pretrained video diffusion models can act as renderers when the desired scene state is already specified by an animated mesh, a camera trajectory, and a reference image. This 4D generative rendering setting raises a representation question: what image-format condition lets a video backbone obey both camera motion and scene-internal animation? We propose DAR, a reference-guided renderer that extends Wan2.2 camera control from Pl\\\"ucker rays alone to a joint camera-plus-geometry interface. DAR projects a neural 4D G-buffer (tracking, world position, and normal) from the animated mesh and injects it through a widened control adapter while preserving the pretrained image-to-video prior. The central design choice is the pair of tracking and world position. Tracking identifies the persistent surface element that should carry appearance; world position gives its current scene-coordinate state; normal supplies local shape. Depth plus calibrated rays can recover 3D in principle, but depth is a camera-dependent chart in which camera and object motion are mixed. On the 68-case DAR-4D benchmark, LoRA DAR reaches PSNR 23.22, SSIM 0.895, and LPIPS 0.134, improving over off-the-shelf Wan2.2-Depth by 1.54 dB PSNR; a full fine-tune reaches PSNR 25.36 and SSIM 0.917. Matched ablations show that replacing world position by depth reduces PSNR by 1.26--1.55 dB at every checkpoint, supporting tracking+world-position correspondence as a practical 4D rendering condition.", "url": "https://wpnews.pro/news/video-models-as-native-4d-renderers-world-grounded-conditioning-from-animated", "canonical_source": "https://arxiv.org/abs/2608.00094", "published_at": "2026-08-04 04:00:00+00:00", "updated_at": "2026-08-04 04:39:08.593152+00:00", "lang": "en", "topics": ["artificial-intelligence", "computer-vision", "generative-ai"], "entities": ["DAR", "Wan2.2", "DAR-4D benchmark"], "alternates": {"html": "https://wpnews.pro/news/video-models-as-native-4d-renderers-world-grounded-conditioning-from-animated", "markdown": "https://wpnews.pro/news/video-models-as-native-4d-renderers-world-grounded-conditioning-from-animated.md", "text": "https://wpnews.pro/news/video-models-as-native-4d-renderers-world-grounded-conditioning-from-animated.txt", "jsonld": "https://wpnews.pro/news/video-models-as-native-4d-renderers-world-grounded-conditioning-from-animated.jsonld"}}