Skip to content

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling

Jun 2026 · arXiv.org · Vol abs/2606.20799 · 3 citations · 64 references
Computer Science

TL;DR

GroundShot is presented, a training-free, model-agnostic agentic framework for entity-grounded multi-shot generation that improves multi-shot consistency over existing methods while requiring no additional training or model modification.

Abstract

Generating visually consistent multi-shot videos remains an open challenge. As videos span more shots, inconsistencies can accumulate across shots, causing entities that reappear across shots -- characters, objects, and locations -- to drift away from how they first appear. We observe that viewers judge consistency by comparing each later appearance of an entity with its first clear appearance; the visual quality of this initial appearance sets the consistency ceiling for all that follows. Motivated by this, we present \textbf{GroundShot}, a training-free, model-agnostic agentic framework for entity-grounded multi-shot generation. GroundShot builds an entity-level visual memory online from accepted generated shots: it schedules shots'generation order by their expected usefulness as entity references, grounds entities from generated videos, verifies their reliability before adding them to memory, and retrieves suitable entity references from memory before each shot is generated. To evaluate this entity-centered view of consistency, we further introduce \textbf{GroundBench}, a diagnostic benchmark that measures consistency at the entity level while isolating controlled challenge dimensions. Experiments show that GroundShot improves multi-shot consistency over existing methods while requiring no additional training or model modification.

View source

Similar papers

Preprint Aug 2026

LogiShot: Logically Coherent Cross-Shot Video Generation

LogiShot is proposed, which incorporates information through two complementary paths that jointly encodes the context video and other conditioning signals, yielding dense multimodal cues that provide visual-semantic evidence for cross-shot generation and the model maintains a visual memory of the context video throughout generation to preserve visual consistency across shots.

Shuai Guo, Yuhang Yang, Zeyu Zhang et al. · 1 citation
Preprint Aug 2026

VideoVIBE: A Video-Grounded Diagnostic Benchmark for One-Shot Interactive Website Generation

Natural-language-driven"vibe coding"enables the one-shot generation of visually rich and interactive web applications, yet reliable assessment of their quality has not kept pace. Existing evaluations often score isolated artifacts or final task outcomes, offering limited evidence about which failures occur and why. We introduce VideoVIBE, a video-grounded benchmark that transforms human-operated webpage recordings into fine-grained diagnostic tasks. It contains approximately 1.7K diagnostic Video QA instances derived from 6,338 verified failures across generated webpages, spanning semantic-logical, visual-motion, structural-temporal, and functional failures. Diagnoses are grounded primarily in recorded presentation and behavior, with webpage source code used as complementary context. We further propose V2Lens, a training-free, evidence-grounded multi-agent system that challenges and selectively refines initial video-based diagnoses through targeted visual and source-code verification. Across thirteen closed-source and open-weight Video MLLMs, Gemini-2.5-Flash is the strongest standalone model with a score of 64.54, while V2Lens reaches 71.72, an improvement of 7.18 points. Together, our results show that video-grounded evaluation can move beyond isolated artifacts and aggregate outcomes toward a behaviorally faithful and diagnostically informative account of generated application quality.

Jiajun Xu, Yanghao Zhou, Jing Liao et al. · 0 citations
#small language model Preprint Aug 2026

SEAM: Shot Entity-Attribute Memory for Consistent Short-Drama Generation at Scale

A training-free, model-agnostic memory graph that repairs continuity entirely at the prompt-text layer by extracting a multi-dimensional state for every shot, retrieving only causally prior context over the resulting graph, filtering it selectively, and injecting the surviving constraints by natural-language prompt rewriting.

Jiaqi Liu, Maolin Ran, Xiaoyan Lu et al. · 0 citations
Jul 2026

Large model-assisted video summarization via global entity unification and robust importance scoring

A summarization pipeline around Global Entity Unification and Robust Importance Scoring is built, but unlike earlier efforts, each object is traced across frames and attached consistent identifiers to it, and Fragmented, isolated descriptions become a single, object-aware text corpus that unifies the storyline.

Donglei Chen, Shaoyu Huang, Xuemiao Xu et al. · 0 citations
Preprint Sep 2026

Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation

Generating long-form multi-shot videos requires coherent within-shot motion and visually consistent narratives across shots. Existing video generators favor continuous motion and struggle to present complete shot sets when an entire narrative is packed along one temporal axis. We propose MovieGrid, a Multi-Grid Post-Training paradigm that decomposes a long video into shorter, temporally ordered chunks and arranges them on a spatial grid for joint modeling. This design reduces the number of shots handled by each temporal axis while enabling global information exchange across chunks. We construct the Multi-Grid Long Video (MGLV) dataset from 1,000 long-form videos using source video collection, hierarchical segmentation, grid video construction, and character-aware story annotation, producing 54K grid videos paired with story prompts. Our Noise-Free Random-Grid Training retains a random subset of chunks as clean visual context for denoising the remaining chunks. Grid Embedding encodes grid structure, character-aware Story Prompts link recurring entities, and Grid Boundary Loss stabilizes layouts. Under the same token budget, MovieGrid generates 6.05 times more shots than Temporal Packing in a 1,616-frame video. On a benchmark spanning five real-world categories, it achieves state-of-the-art intra-shot consistency (0.9131 versus 0.8086 for HoloCine) and inter-shot consistency (0.5914 versus 0.5384 for StoryMem). MovieGrid can further scale video length with minimal compromise through single or multiple generations.

Unknown authors · 0 citations
Preprint Aug 2026

MSEditor: Toward Consistent Multi-Shot Video Editing

MSEditor is proposed, the first framework designed specifically for consistent multi-shot video editing, which significantly outperforms existing methods on the authors' curated multi-shot video editing benchmark in terms of identity preservation, temporal stability, and overall visual quality.

Kunyu Feng, Yue Ma, Bingyuan Wang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.