Jul 2026· Journal of Cybersecurity and Privacy· Vol 6, pp. 121· 0 citations· 35 references
TL;DR
This work proposes Text-to-Unlearn, a novel framework that selectively unlearns concepts from pre-trained GANs using only text prompts, enabling feature and identity unlearning, as well as fine-grained tasks such as expression and multi-attribute removal in models trained on human faces.
Abstract
State-of-the-art generative models exhibit powerful image-generation capabilities, raising ethical and legal challenges for service providers. Consequently, Content Removal Techniques (CRTs) have emerged to control outputs without requiring full retraining. However, the problem of unlearning in Generative Adversarial Networks (GANs) remains largely unexplored. We propose Text-to-Unlearn, a novel framework that selectively unlearns concepts from pre-trained GANs using only text prompts, enabling feature and identity unlearning, as well as fine-grained tasks such as expression and multi-attribute removal in models trained on human faces. Our approach leverages natural language descriptions to guide unlearning without additional datasets or supervised finetuning, offering a scalable solution. To evaluate the effectiveness of our method, we introduce an automated unlearning assessment method using state-of-the-art image–text alignment metrics and propose a new metric: degree of unlearning. Additionally, we assess robustness by introducing adversarial attacks to subvert unlearning. Our results demonstrate that Text-to-Unlearn achieves robust unlearning, resisting adversarial attempts to recover erased concepts while preserving model utility. To our knowledge, this is the first cross-modal unlearning framework for GANs, advancing the management of generative model behavior.
The results reveal that attention-space regulation offers a considerably more promising path to safer diffusion transformer based image generation than the existing concept erasing mechanism.
Text-to-image generation is an increasingly fast-paced field of generative artificial intelligence, consisting of synthesizing images of high quality and semantic consistency based on natural language descriptions. In this paper, we give an extensive overview of the approach to text-to-image generation using deep learning, including the most common core model families, architecture designs, training approaches, and evaluation systems. We discuss the paradigms of the generative adversarial networks (GANs), variational autoencoders (VAEs), transformer-based designs, and diffusion models, with the last one representing the state of the art in image generation models. The review also discusses key aspects of pipelines such as text encoding, cross-modal alignment, mechanisms of attention, and decoding images. Popular datasets, methods, and metrics of evaluation, including Fréchet Inception Distance (FID) and CLIP-based similarity, are discussed. The application domains that involve creative content creation, medical imaging, education and industrial design are critically discussed. Despite significant advances, various issues still exist, such as low stability in training, excessive computational complexity, amplification of bias, generated images, and text–image alignment errors. Moral and social issues, such as misinformation, intellectual property, and equity, are critically examined. Lastly, we present future research directions to more controllable, more efficient and more interpretable text-to-image systems, focusing on multimodal foundation models and human–AI collaborative design.
DiSCO is proposed, a zero-shot, strictly black-box defense that operates entirely at the prompt level as a plug-and-play module, requiring no model retraining, fine-tuning, or access to model internals, and can be readily applied to any text-to-image system without necessitating any changes to the model itself.
Tong Zhang, M. Alfarra, Carlos Hinojosa et al.· 0 citations
A lightweight Text Encoder Alignment framework that fine-tunes only the text encoder while keeping the generative backbone fully frozen, and achieves state-of-the-art erasure robustness against black-box and white-box adversarial attacks on Stable Diffusion v1.4, while preserving generation quality on benign prompts.
Large Vision-Language Models (LVLMs) have demonstrated remarkable multimodal comprehension capabilities, achieving state-of-the-art performance across various vision-language tasks. However, their performance drops significantly when facing adversarial attacks on the visual encoder. To alleviate this issue, existing approaches often rely on adversarial training, enhancing model robustness through substantial computational cost. Unlike these methods, this paper proposes a novel, training-free adversarial defense method called Edge-Guided Prompt Defense (EGP-Defense), which performs adversarial defense during the model inference stage. This method is based on a comprehensive analysis of image edges under various types of attacks. We observe that edge maps exhibit strong robustness against adversarial attacks, and the extracted edge features can effectively reflect key aspects of the original image. Building on this observation, we first apply the Canny operator to extract edge maps from input images, and then use LVLMs to generate textual descriptions based on these structural representations. To further distill the most critical information from these descriptions, we extract informative keywords and incorporate them as auxiliary prompts. These prompts guide the model to focus on task-relevant features during inference, thereby enhancing its robustness against adversarial perturbations. Extensive experiments demonstrate that EGP-Defense significantly improves the robustness of LVLMs against three types of adversarial attacks in both image classification and image caption tasks.
Bo-Yu Wang, Zi-Wen He, Xin-Jue Hu et al.· ACM Transactions on Multimed...· 0 citations
This work introduces PRMU, a benchmark for evaluating corpus-free multimodal unlearning under realistic person-centric deletion requests, and introduces Similarity-Gated Projection Editing (SGPE), a lightweight corpus-free unlearning baseline with knowledge displacement, protected parameter-space editing, and locality-aware multimodal control.
Hua-Feng Chen, Yueming Lyu, Ziyuan Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.