Google Research details four agentic frameworks for coherent long-form video generation
Original titleAutomating coherent long-form video generation
Google Research introduces four multi-agent frameworks for generating minutes-long videos with consistent characters and environments across shots.
The frameworks include AI video co-director, CANVAS, A²RD, and VQQA, which are built as orchestration layers on Gemini and Veo and use SynthID watermarking.
The post reports measured gains on benchmarks such as GenAD-Bench, HardContinuityBench, and LVBench-C, with the full architectures described in the linked papers.
The post links four frameworks to specific failure modes in long video generation, such as semantic drift and cascading errors, making the design choices easier to compare.
Source: Google Research · research.googlePublished · added here