Google unveils AI video co-director with four frameworks for seamless long-form video creation
Google Research has unveiled an AI video co-director, leveraging four agentic frameworks to generate coherent, minutes-long videos. The system addresses common issues like identity drift and cascading errors in multi-shot video pipelines. It operates on top of Gemini and Veo, with outputs inheriting SynthID watermarking. The framework, accepted at COLM 2026, uses a multi-armed bandit approach with specialized agents for creative strategy, narrative mode, and aesthetic archetype.
CANVAS, accepted at EMNLP 2026, maintains consistency in characters, locations, and object states, as demonstrated in a museum heist test. A²RD, a training-free architecture, generates a 10-minute film using a Retrieve-Synthesize-Refine-Update loop. VQQA, which generates visual questions for prompts, uses a Global Selection step to choose the best video output.
Google also introduced three new benchmarks: GenAD-Bench, HardContinuityBench, and LVBench-C, each testing different aspects of video continuity and coherence. The project’s details are available through linked papers, project pages, and GitHub repositories, verified on September 27, 2026.