Google Research Introduces an AI Video Co-Director: 4 Agentic Frameworks for Coherent, Minutes-Long Video Generation Google Research introduced an AI video co-director suite of four agentic frameworks that turns short clips into coherent, minutes-long stories by targeting identity drift and cascading errors in multi-shot AI video pipelines. The system, built on Gemini and Veo and described as model-agnostic, includes Co-Director, CANVAS, A²RD, and VQQA; Co-Director scored 81.4 average on the new GenAD-Bench and 3.96 of 5 in human ratings, while CANVAS delivered gains of 21.6% in background continuity, 9.6% in character consistency, and 7.6% in props consistency, and A²RD improved consistency by up to 30% and narrative coherence by 20% on 1 to 10 minute videos. Google also built three benchmarks: GenAD-Bench with 400 ad scenarios across 200 fictional products from 50 brands, HardContinuityBench, and LVBench-C with 120 scenarios where key assets vanish for at least 10 segments before returning. Google Research https://research.google/blog/coherent-long-form-video-generation/ has introduced an AI video co-director for long-form video generation. The suite of 4 agentic frameworks turns short clips into coherent, minutes-long stories. It targets identity drift and cascading errors, the 2 failures that break most multi-shot AI video pipelines today. Why Long AI Videos Fall Apart Diffusion models render high-fidelity clips in seconds. Stitching those clips into a story is harder. Most agentic pipelines chain modules with independent, handcrafted prompts. That causes semantic drift https://arxiv.org/abs/2511.17986 , where attire or scenery shifts between shots. It also causes cascading failures https://arxiv.org/abs/2606.24976 , where one bad upstream asset corrupts every later shot. Google team frames this as a credit assignment problem. A broken final video is hard to trace back to the prompt that caused it. How the AI Video Co-Director Works The system sits on top of Gemini https://deepmind.google/models/gemini/pro/ and Veo. It is model-agnostic, so the same layer can drive other generators. Outputs inherit SynthID https://deepmind.google/models/synthid/ watermarking from the base models. 1. Co-Director: creative planning as a bandit search Co-Director https://arxiv.org/abs/2604.24842 , accepted at COLM 2026, uses a multi-armed bandit MAB . An Orchestrator Agent picks a configuration across Creative Strategy, Narrative Mode, and Aesthetic Archetype. A Pre-Production Agent builds the storyboard. Keyframe, Video, and Audio sub-agents produce the media. An MLLM Judge then scores the cut and sends a factored reward back to the bandit. 2. CANVAS: persistent visual memory CANVAS https://arxiv.org/abs/2604.13452 , accepted at EMNLP 2026, tracks characters, locations, and object states as the story evolves. It retrieves stored visual anchors when a scene returns. In Google’s museum heist test, AutoStudio https://arxiv.org/abs/2406.01388 lost the thief’s cap and Gemini-3.1-Pro changed the gemstone. CANVAS kept both consistent. 3. A²RD: segment-by-segment long video A²RD https://arxiv.org/abs/2605.06924 Agentic Autoregressive Diffusion is a training-free architecture. Each segment runs a Retrieve, Synthesize, Refine, Update loop against a multimodal video memory. The agent switches between extrapolation for new story beats and interpolation for returning entities. Google shared a 10-minute film https://www.youtube.com/watch?v=jaLZK8Zb6Kc generated this way. 4. VQQA: closed-loop prompt refinement VQQA https://arxiv.org/abs/2603.12310 Video Quality Question Answering generates visual questions for each prompt. VLM critiques act as “semantic gradients” that rewrite the text prompt. It needs no access to model internals. A Global Selection step picks the best video across all iterations, not simply the last one. Benchmarks and Results Google built 3 new benchmarks. GenAD-Bench https://co-director-agent.github.io/genad bench.html has 400 ad scenarios across 200 fictional products from 50 brands. HardContinuityBench stresses scene reappearances and prop state changes. LVBench-C https://github.com/dxlong2000/AARD has 120 scenarios where key assets vanish for at least 10 segments before returning. - Co-Director: 81.4 average on GenAD-Bench and 3.96 of 5 in human ratings, per the project page https://co-director-agent.github.io/ . Baselines included Veo 3.1, Kling 3.0 Omni, Wan 2.6, and MovieAgent. - CANVAS: gains of 21.6% in background continuity, 9.6% in character consistency, and 7.6% in props consistency. - A²RD: up to 30% better consistency and 20% better narrative coherence on 1 to 10 minute videos. - VQQA: absolute gains of 11.57% on T2V-CompBench https://github.com/KaiyueSun98/T2V-CompBench and 8.43% on VBench2 https://github.com/Vchitect/VBench over vanilla generation. How It Compares | Feature | Google AI Video Co-Director https://research.google/blog/coherent-long-form-video-generation/ | StoryMem https://kevin-thu.github.io/StoryMem/ ByteDance, NTU | MovieAgent https://arxiv.org/abs/2503.07314 Show Lab, NUS | AutoStudio https://arxiv.org/abs/2406.01388 | |---|---|---|---|---| | Output | Minutes-long multi-shot video with voiceover and score | Minute-long multi-shot video | Multi-scene, multi-shot video with subtitles and audio | Multi-turn image sequences no video | | Architecture | 4 frameworks in a hierarchical multi-agent orchestration layer | Memory-to-Video diffusion model, shot by shot | Multi-agent chain-of-thought planning director, screenwriter, storyboard artist, location manager | 3 LLM agents plus a Stable Diffusion based agent | | Consistency mechanism | Persistent visual memory CANVAS and multimodal video memory A²RD | Keyframe memory bank from earlier shots | Hierarchical planning plus per-character customization | Subject manager plus Parallel-UNet | | Self-correction loop | Bandit search with MLLM Judge; VQQA prompt refinement with Global Selection | Semantic keyframe selection and aesthetic filtering | Not reported | Not reported | | Model training | No fine-tuning; orchestrates existing models | LoRA fine-tuning on the base model | Per-character LoRA ED-LoRA via ROICtrl | Training-free | | Base generators | Gemini and Veo model-agnostic | Wan2.2 | ROICtrl, SVD, HunyuanVideo I2V | Stable Diffusion | | Longest reported output | 10 minutes A²RD | About 1 minute | Not specified | N/A images | | Code | Co-Director https://github.com/GoogleCloudPlatform/genmedia-izumi-agent/tree/main/demos/backend/ads codirector and A²RD https://github.com/dxlong2000/AARD public; CANVAS coming soon | Public https://github.com/Kevin-thu/StoryMem | Public https://github.com/showlab/MovieAgent | Public https://github.com/donahowe/AutoStudio | Sources: linked papers, project pages, and GitHub repositories. Verified September 27, 2026. Key Takeaways - Google treats long-form video as a global optimization and world-state tracking problem. - 4 frameworks cover planning, storyboarding, long generation, and self-correction. - It runs as an orchestration layer on Gemini and Veo, with SynthID watermarking. - A²RD produced a continuous 10-minute film with stable characters and locations. - Co-Director scored 81.4 on GenAD-Bench, ahead of a 75.7 random search baseline. FAQ - What is Google’s AI video co-director? It is a multi-agent orchestration layer on Gemini and Veo. It plans, generates, and corrects multi-shot videos to keep them consistent. - How long can the videos be? A²RD was evaluated on videos from 1 to 10 minutes, and Google released a continuous 10-minute demo film. - Can developers use it today? Partially. Co-Director and A²RD code is on GitHub. CANVAS code is pending, and the full pipeline is not a Google product. Check out the Technical Details https://research.google/blog/coherent-long-form-video-generation/? . All credit goes to the researcher of this project. Also, feel free to follow us on Twitter https://x.com/intent/follow?screen name=marktechpost and don’t forget to join our 150k+ML SubReddit https://www.reddit.com/r/machinelearningnews/ and Subscribe to our Newsletter https://magic.beehiiv.com/v1/f5e63dd4-5653-4f09-83e2-321a8b1ba526?email={{email}} . Wait are you on telegram? now you can join us on telegram as well. https://t.me/machinelearningresearchnews Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us https://forms.gle/MJjjVDPS7whH8Ngs6 Asif Razzaq is the CEO of Marktechpost AI Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.