Skip to content
Conference

AI-Based Personalized Movie Scene Reimagining System for Text-Guided Video Synthesis

Aug 2026 · International Conference on Circuit, Power and Computing Technologies · pp. 1654-1660 · 0 citations · 31 references

Abstract

The generative artificial intelligence has made tremendous advancement in visual content generation; but the issue of intuitive and user-friendly adjustment of the current video scenes is a difficult question to answer. The paper describes a personalized movie reimagining system based on AI that reimagines input video scenes based on natural language instructions. In contrast to the traditional text-to-video methods, which make the content anew, the suggested structure of the visual content is based on structure-preserving transformation, meaning that users have the opportunity to adjust visual properties while preserving the original composition of the scene and motion dynamics. It combines computer vision to comprehend the scene, natural language processing to decipher user intent, and diffusion-based generative models to synthesize videos. ControlNet using Canny edge conditioning is used to maintain spatial structure and motion adapters in AnimateDiff maintain temporal coherence across frames. Also, a proactive enhancement mechanism and dynamic conditioning plan enhance congruence between user input and output generated. The results of the experimental work conducted on a variety of scene types indicate that the proposed system reaches a semantic alignment score of 4.2/5 and structural similarity index (SSIM) of 0.81, which is better than the baseline text-to-video-based methods. The system produces the short video sequences (8-24 frames) in 120-180 seconds with the standard GPU hardware. These results indicate the usefulness of the framework in facilitating high-quality video transformation.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.