Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion models. Yet video understanding benchmarks still evaluate models mainly through question answering. If an agent truly understands a video, it can reconstruct it programmatically. We introdu
LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents