Jul 2026· American Journal of AI Cyber Computing Management· Vol 6, pp. 391-396· 0 citations· 4 references
TL;DR
A semantic-aware framework for text-to-face image synthesis using a joint Bidirectional Long Short-Term Memory (BiLSTM) network and Generative Adversarial Network (GAN) that improves semantic consistency and visual realism.
Abstract
Recent advancements in deep learning have enabled the generation of realistic images directly from natural language descriptions. This paper presents a semantic-aware framework for text-to-face image synthesis using a joint Bidirectional Long Short-Term Memory (BiLSTM) network and Generative Adversarial Network (GAN). The proposed approach simultaneously trains the text encoder and image generator, allowing effective learning of semantic relationships between textual attributes and facial features. Initially, input descriptions are transformed into meaningful vector representations using Bi-LSTM, which are then utilized by the GAN to synthesize high-quality facial images. Unlike conventional methods that rely on separately trained text encoders, the proposed end-to-end architecture improves semantic consistency and visual realism. The model is trained on the CelebA dataset with corresponding facial descriptions and evaluated using similarity and image quality measures. Experimental results demonstrate improved face generation accuracy and better preservation of facial attributes, making the framework suitable for applications in forensic investigations, digital character creation, intelligent human-computer interaction, and public safety systems.
Text-to-image generation is an increasingly fast-paced field of generative artificial intelligence, consisting of synthesizing images of high quality and semantic consistency based on natural language descriptions. In this paper, we give an extensive overview of the approach to text-to-image generation using deep learning, including the most common core model families, architecture designs, training approaches, and evaluation systems. We discuss the paradigms of the generative adversarial networks (GANs), variational autoencoders (VAEs), transformer-based designs, and diffusion models, with the last one representing the state of the art in image generation models. The review also discusses key aspects of pipelines such as text encoding, cross-modal alignment, mechanisms of attention, and decoding images. Popular datasets, methods, and metrics of evaluation, including Fréchet Inception Distance (FID) and CLIP-based similarity, are discussed. The application domains that involve creative content creation, medical imaging, education and industrial design are critically discussed. Despite significant advances, various issues still exist, such as low stability in training, excessive computational complexity, amplification of bias, generated images, and text–image alignment errors. Moral and social issues, such as misinformation, intellectual property, and equity, are critically examined. Lastly, we present future research directions to more controllable, more efficient and more interpretable text-to-image systems, focusing on multimodal foundation models and human–AI collaborative design.
The results indicate that developing domain-aware alignment and hybrid loss integration techniques is beneficial for effective facial analysis in both controlled and challenging environments.
Huihui Yin, Yurui Guan· International Conference on...· 0 citations
This study builds and test a Deep Convolutional Generative Adversarial Network (DCGAN) that can produce realistic portraits of people's faces and proves that DCGANs are capable of creating realistic facial representations.
K. N. Reddy, A. Renuka· International Journal for Re...· 0 citations
A diffusion-based framework with a stage-wise multi-condition guidance mechanism that enhances both structural and textural fidelity and compares with recent image-to-image translation and diffusion-based baselines to observe competitive performance in both visual coherence and identity preservation.
Yue Que, Xuegui Cheng, Shuqian Shi et al.· Neural Networks· 0 citations
Facial expression recognition technology is vital for security, verification, and personalization, but it faces challenges due to variations in scale, illumination, occlusion, and facial expressions. This paper presents a hybrid architecture that combines Vision Transformers (ViTs) to capture global context with EfficientNet-B3 for multi-scale feature extraction. Unlike simple concatenation, our approach projects the ViT’s [CLS] token and the EfficientNet’s global pooling features into a shared 512-dimensional space before merging, enabling better alignment of global and local features. When tested on the FERPlus dataset, it reaches an accuracy of 94.4 ± 0.3%, surpassing several recent methods, notably existing transformer- and CNN-based methods. Ablation studies show each component’s contribution, with the full model outperforming the no-fusion version by 2.6%. With around 98 million parameters and an inference time of ~23 ms per image, it balances efficiency and high performance, suitable for real-time use on suitable hardware. Evaluation via confusion matrix, t-SNE visualization, and comparisons with recent techniques such as HLA-ViT (90.13%), AU-ViT (90.15%), and CCFER (91.24%) demonstrates its robustness and discriminative feature learning. This work highlights the promise of hybrid deep learning architectures in tackling real-world facial expression recognition challenges.
Sasan Karamizadeh, Saman Shojae Chaeikar, Mazdak Zamani· Journal of Imaging· 0 citations
The rising demand for facial attribute editing has highlighted the limitations of existing text-driven generation methods. Although these methods can enhance feature diversity, they typically depend on large-scale annotated datasets, with varying text guidance necessitating separate optimization processes. Additionally, pre-trained Vision-Language Models (VLMs) have limitations in extracting text information and tuning hyperparameters. To tackle these challenges, this paper introduces a facial attribute editing method utilizing Dual-Branch Mapping (DBM). The core concept of this method is to achieve feature fusion between the source and target images in vector space. This method calculates the vector difference between the source and target images in the Contrastive Language-Image Pre-training (CLIP) space using both global and local branches. The difference is then fused into the Style Generative Adversarial Network (StyleGAN) vector of the source image, enabling precise control over attribute changes. During the inference stage, predictions and adjustments are made based on the correlation matrix in the StyleGAN space, following the change direction indicated by the CLIP text space. Experimental results indicate that this method enables text-free guided facial attribute editing, effectively overcoming the limitations of parameter tuning and prior data. Furthermore, it supports image generation under multiple text conditions without requiring additional training, significantly enhancing editing flexibility and naturalness.
Hu He, Jiayang Yu, Guanghua Gu et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.