Skip to content
Open access

AI-Based Text-to-Video Generation Using Diffusion and Deep Learning Techniques

Jul 2026 · International Journal for Research in Applied Science and Engineering Technology · Vol 14, pp. 634-640 · 0 citations

TL;DR

An AI-based Text-to-Video Generation framework that converts natural language prompts into short animated video sequences using diffusion and deep learning techniques is presented, providing a lightweight and costeffective platform for video generation.

Abstract

Recent advances in Generative Artificial Intelligence have significantly improved automatic multimedia content creation. This paper presents an AI-based Text-to-Video Generation framework that converts natural language prompts into short animated video sequences using diffusion and deep learning techniques. The proposed framework employs the Stable Diffusion v1.5 model for generating high-quality image frames and AnimateDiff for introducing smooth temporal motion while preserving scene consistency. The DDIM scheduler is utilized to accelerate the denoising process and improve inference efficiency without compromising output quality. The generated frames are sequentially processed and assembled into MP4 videos using FFmpeg. The proposed system is implemented using Python, PyTorch, Hugging Face Diffusers, and Google Colab, providing a lightweight and costeffective platform for video generation. Experimental evaluation demonstrates that the proposed framework generates visually realistic and semantically meaningful videos with improved motion continuity and reduced computational complexity compared to conventional frame-by-frame generation approaches. The proposed framework has potential applications in digital media, education, entertainment, advertising, animation, and content creation

Read PDF

Similar papers

Conference Aug 2026

AI-Based Personalized Movie Scene Reimagining System for Text-Guided Video Synthesis

The generative artificial intelligence has made tremendous advancement in visual content generation; but the issue of intuitive and user-friendly adjustment of the current video scenes is a difficult question to answer. The paper describes a personalized movie reimagining system based on AI that reimagines input video scenes based on natural language instructions. In contrast to the traditional text-to-video methods, which make the content anew, the suggested structure of the visual content is based on structure-preserving transformation, meaning that users have the opportunity to adjust visual properties while preserving the original composition of the scene and motion dynamics. It combines computer vision to comprehend the scene, natural language processing to decipher user intent, and diffusion-based generative models to synthesize videos. ControlNet using Canny edge conditioning is used to maintain spatial structure and motion adapters in AnimateDiff maintain temporal coherence across frames. Also, a proactive enhancement mechanism and dynamic conditioning plan enhance congruence between user input and output generated. The results of the experimental work conducted on a variety of scene types indicate that the proposed system reaches a semantic alignment score of 4.2/5 and structural similarity index (SSIM) of 0.81, which is better than the baseline text-to-video-based methods. The system produces the short video sequences (8-24 frames) in 120-180 seconds with the standard GPU hardware. These results indicate the usefulness of the framework in facilitating high-quality video transformation.

Thupakula Venkata Sumanth, Kriti Gupta, Charanjit Singh et al. · 0 citations
Review Open access Jul 2026

Text-to-Image Generation via Deep Learning: A Comprehensive Review of Models, Architectures, and Future Directions

Text-to-image generation is an increasingly fast-paced field of generative artificial intelligence, consisting of synthesizing images of high quality and semantic consistency based on natural language descriptions. In this paper, we give an extensive overview of the approach to text-to-image generation using deep learning, including the most common core model families, architecture designs, training approaches, and evaluation systems. We discuss the paradigms of the generative adversarial networks (GANs), variational autoencoders (VAEs), transformer-based designs, and diffusion models, with the last one representing the state of the art in image generation models. The review also discusses key aspects of pipelines such as text encoding, cross-modal alignment, mechanisms of attention, and decoding images. Popular datasets, methods, and metrics of evaluation, including Fréchet Inception Distance (FID) and CLIP-based similarity, are discussed. The application domains that involve creative content creation, medical imaging, education and industrial design are critically discussed. Despite significant advances, various issues still exist, such as low stability in training, excessive computational complexity, amplification of bias, generated images, and text–image alignment errors. Moral and social issues, such as misinformation, intellectual property, and equity, are critically examined. Lastly, we present future research directions to more controllable, more efficient and more interpretable text-to-image systems, focusing on multimodal foundation models and human–AI collaborative design.

Abdussalam Elhanashi, Siham Essahraui, Qinghe Zheng et al. · 0 citations
Jul 2026

Enhanced Image Captioning Using Sinusoidal Spatial Transform CNN and Fine-tuned BERT

Image caption generation is a significant area of study in artificial intelligence and computer vision, focusing on training systems to generate accurate textual descriptions of images. This paper presents an advanced framework for image captioning using a dataset sourced from Kaggle. The process begins with pre-processing techniques, including the Flexiframe Filter, which dynamically adjusts the window size based on local variance to reduce noise in both text and images. Bright- Contrast augmentation is then applied to enhance input images, enriching the dataset and improving feature extraction. The encoder module utilises a Sinusoidal Spatial Transform Convolutional Neural Network (SST-CNN) integrated with a Visual Geometry Group (VGG16) model for feature extraction and encoding. For the decoder, stochastic k-sampling is combined with a Dense Neural Network and a Gated Recurrent Unit (GRU) to generate initial captions. To refine these captions, a fine-tuned Bidirectional Encoder Representations from Transformers (BERT) model is employed for enhanced coherence and accuracy. The model's performance is evaluated using standard metrics, achieving BLEU-1, METEOR, ROUGE-L, and CIDER scores of 1, 0.99, 0.99, and 1, respectively. The proposed approach demonstrates significant improvements in generating descriptive and accurate image captions, making it a robust solution for applications requiring semantic understanding of visual data.

Bhargavi Polepalli, Praveen Kumar Sekharamantry, K. S. Rao · 0 citations
Open access Jul 2026

Attention-Based Deep Learning Pipeline for AI-Created Image Recognition

The proposed Attention-Based Deep Learning Pipeline of AI-Created Image Recognition incorporates three integrated branches, including low-level statistical feature extraction, high-level semantic representation learning, and attention-based feature refinement mechanism, which support the robustness and generalization ability of the proposed model in detecting AI-generated images in a variety of generators and conditions.

Nadia Ali · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.