Skip to content
Conference

Scene-Level Planning for Temporally Coherent AI Video Generation Using Multimodal Representations

Aug 2026 · 2026 International Conference on Intelligent Multimedia, Networking, and Security (IMNS) · pp. 1-6 · 0 citations · 12 references

Abstract

Recent text-to-video systems can generate visually appealing clips from natural language prompts, yet narrative prompts often contain multiple implicit temporal stages that require the generator to infer scene decomposition, subject persistence, action ordering, and visual continuity from a single unstructured input. This frequently leads to temporally inconsistent or structurally ambiguous outputs. In this work, we investigate whether introducing an explicit scene-planning layer can improve multi-stage video generation. We compare three generation paradigms: direct single-prompt generation, naive prompt decomposition, and a structured scene-planning pipeline. The proposed approach first converts a narrative prompt into a lightweight structured scene representation containing global subject information, visual style constraints, and scene-level descriptions, which is then compiled into scene-conditioned prompts for sequential video generation. A reference-guided continuation mechanism conditions the second scene on the final frame of the first scene to improve cross-scene identity and visual continuity. The framework is generator-agnostic and can operate on modern text-to-video backends without modifying the underlying models. To evaluate generation quality, we adopt a vision-language model (VLM) as an automatic judge that assesses prompt relevance, temporal continuity, aesthetic quality, and narrative clarity across candidate videos. This study provides an empirical investigation of how structured intermediate planning influences narrative video generation and offers a lightweight framework for improving temporal coherence in AI-generated videos.

View source

Similar papers

Jul 2026

VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation

Experiments on an unseen validation set show that VIPER achieves stronger reference-video physical similarity and higher human preference than representative video generation and video-as-prompt baselines, while maintaining competitive general video quality.

Tianxi Chen, Han-Mo Chen, Hua-Jin Chen et al. · 0 citations
Jul 2026

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

VideoCoCo, an agentic dual-engine framework in which executable Blender code serves as a process-level chain of thought, demonstrates that executable code provides an effective, controllable, and inspectable intermediate representation for physically consistent video generation.

Haodong Li, Tianfei Ren, Xiaoxiao Ma et al. · 8 citations
Preprint Aug 2026

CoT-Edit: Let CoT Guide Instruction Video Editing

Text-driven instruction-based video editing in complex scenes remains challenging: purely textual prompts often fail to capture precise spatial relationships and physical constraints, resulting in target ambiguity and physically implausible outcomes. To address this, we propose a plan--guide--edit framework that explicitly bridges semantic intent and spatial execution. In our framework, a Chain-of-Thought (CoT)-enhanced multimodal large language model (MLLM) serves as a planner, performing structured reasoning over the video and instructions to derive a precise sequence of bounding boxes and attribute-enriched editing directives. These spatial priors then guide a box-conditioned mask generator, transforming ambiguous global retrieval into localized, context-aware refinement and producing masks that more accurately capture object scale, contact relationships, and placement. Building on these spatial and semantic signals, a diffusion-based editor integrates the masks, enriched instructions, and frame features to render high-fidelity edits that remain temporally coherent and spatially well aligned. Trained first in a modular manner and then jointly, our framework achieves superior performance with reduced data requirements, delivering precise localization in scenes with multiple similar objects and physically consistent object additions, and extensive experiments demonstrate state-of-the-art performance over multiple strong baseline methods. More details are available at: https://github.com/flying-sky999/CoT-Edit

Sen Liang, Fengbin Guan, Youliang Zhang et al. · 5 citations
Open access Aug 2026

Long-Horizon Video Generation with Temporally Consistent Diffusion and Scene-Graph Guidance

A novel framework that integrates temporally consistent diffusion models with dynamic scene-graph guidance that structurally constrains the generative process, ensuring that objects, their attributes, and their interrelationships remain stable over extended durations is introduced.

Jacob A. Jenkins · 0 citations
Preprint Aug 2026

LogiShot: Logically Coherent Cross-Shot Video Generation

LogiShot is proposed, which incorporates information through two complementary paths that jointly encodes the context video and other conditioning signals, yielding dense multimodal cues that provide visual-semantic evidence for cross-shot generation and the model maintains a visual memory of the context video throughout generation to preserve visual consistency across shots.

Shuai Guo, Yuhang Yang, Zeyu Zhang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.