Author

Jacob A. Jenkins

1 paper indexed here

Fetches their full publication history.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Long-Horizon Video Generation with Temporally Consistent Diffusion and Scene-Graph Guidance

The synthesis of high-fidelity, temporally coherent long-horizon videos remains a profound challenge in the domain of generative artificial intelligence. Current diffusion-based approaches often suffer from severe temporal degradation, semantic drift, and structural inconsistency when generating sequences beyond a few seconds. To address these limitations, this paper introduces a novel framework that integrates temporally consistent diffusion models with dynamic scene-graph guidance. By leveraging scene graphs as explicit semantic anchors across frames, the proposed architecture structurally constrains the generative process, ensuring that objects, their attributes, and their interrelationships remain stable over extended durations. The methodology involves a dual-stream architecture where a graph neural network processes sequential scene graphs to condition a cascaded video diffusion model. Furthermore, a specialized spatiotemporal cross-attention mechanism is developed to align latent noise representations with the relational data embedded in the scene graphs. Extensive empirical evaluations on standard video generation benchmarks demonstrate that the proposed method significantly outperforms baseline approaches in both quantitative metrics and qualitative human assessments, particularly in maintaining entity persistence and logical action progression over long time horizons. The findings underscore the critical role of explicit structural representations in overcoming the inherent memory limitations of pure attention-based video generation systems

Jacob A. Jenkins · 0 citations