A two-stage framework for generating privacy-preserving synthetic eye images based on a dataset of real photographs, inspired by the DreamBooth method, which suggests the framework is a promising solution for anonymising biometric eye images in clinical and research applications.
Abstract
Advances in realistic image synthesis have led to several downstream applications in healthcare, particularly the anonymisation of biometric images. However, generative models pose significant privacy risks when synthetic outputs closely resemble the real, personal images required for training. To address these challenges, we propose a two-stage framework for generating privacy-preserving synthetic eye images based on a dataset of real photographs. Our system consists of a fine-tuned text-to-image synthesiser based on Stable Diffusion (v1-5), followed by a privacy risk scorer termed the “Cone of Privacy” (CoP). Inspired by the DreamBooth method, the synthesiser incorporates information about the original dataset to generate similar realistic yet diverse images. To mitigate identity leakage, the CoP score measures the identifiability of synthetic images with respect to the training images based on embeddings from a pre-trained vision transformer. In evaluations on a dataset of 2,000 images from 704 subjects, our approach successfully generated realistic images while filtering high-risk samples, achieving a privacy threat detection rate above 97%. The Fréchet Inception Distance (FID) between the original and the fully synthetic datasets was 57.1, with privacy-violating synthetic images scoring 51.3 and privacy-preserving images scoring 82.6, demonstrating a trade-off between realism and privacy protection. These results suggest the framework is a promising solution for anonymising biometric eye images in clinical and research applications.
A comparative study of four audits applicable to pre-trained, black-box face generators, which consistently reveal substantial identity distinguishability while reporting markedly different epsilon estimates that reflect each method's distinct assumptions and finite-sample treatment.
Arman Zareian Jahromi, Vishnu Bondalakunta, M. Shah et al.· 0 citations
The rapid advancement of sophisticated generative models has intensified the need for robust fake image detection systems. However, many existing benchmark datasets suffer from limited diversity in content types and generation techniques, constraining the generalization ability of detection models. To address these limitations, we introduce GenPix (Generalized Pixels), a comprehensive dataset encompassing over 80,000 images spanning diverse categories, including faces, objects, and scenes, generated by multiple state-of-the-art models such as Generative Adversarial Networks (GANs) and diffusion-based architectures. The dataset includes samples from different generation methods to ensure broad coverage of fake image characteristics.
GenPix provides a realistic evaluation environment that better reflects real-world detection challenges. We establish baseline performance metrics using an Adversarial Autoencoder (AAE) and demonstrate the dataset's utility for developing and evaluating fake image detection systems. The AAE achieves 80.65% F1-score on the full GenPix test set and high inference throughput (488 images/sec).These results show that even relatively simple architectures can achieve promising performance on GenPix, while highlighting areas for improvement in detection methodologies.In contrast, deeper CNNs such as EfficientNet-B3 reach higher F1-score of 98.01% but suffer from low throughput (14 images/sec), suggesting a complementary trade-off between performance and practicality. Overall, GenPix provides a challenging and realistic benchmark for evaluating modern detectors, and the proposed AAE offers an efficient, interpretable baseline for future research on general-purpose fake-image detection.
Guessoum Dalila, B. Nadjia, Boumahdi Fatima et al.· Iraqi Journal for Computer S...· 0 citations
Vein recognition is a secure biometric technology often constrained by limited annotated data and imaging variations. While data augmentation mitigates this, strategies designed for natural images may disrupt the fine-grained topology and textures essential for identity discrimination. We present AGVBench, which evaluates 30 representative augmentation strategies on five public palm- and finger-vein datasets with seven backbone architectures, covering classic CNNs, vision transformers, and vein-specific recognition models. Our results show that multi-image mixing methods (e.g., MixUp, PuzzleMix, StarMixup) generally provide the strongest recognition performance. However, they are often poorly calibrated and vulnerable to adversarial perturbations, revealing a clear inconsistency between clean accuracy and adversarial security. We also find that severe geometric transformations frequently degrade recognition, which is potentially due to feature misalignment or spatial cropping, and that augmentation effectiveness varies across palm and finger vein datasets. These findings prove that accuracy-centric evaluation is insufficient for biometric augmentation. AGVBench provides standardized protocols to support reproducible research and guide the design of reliable, secure, and robust vein recognition systems. Our codebase is available at https://github.com/Advance-VeinTech-Innovators/AGVBench.
Haiyang Li, Yuming Fu, Qun Song et al.· 0 citations
Steganography plays a critical role in preserving the privacy of medical images by protecting sensitive patient information without compromising image quality or diagnostic utility. This paper proposes a novel lossless privacy protection algorithm for medical images based on pixel displacement. The core innovation lies in the synergistic combination of physical transformation and deep learning—specifically, the integration of shift-rectification with a dual-carrier embedding mechanism. The proposed method begins by applying byte-level pixel shifting to the confidential medical image to generate a shift-rectified version. A deep convolutional neural network is then employed to embed both the original and the rectified medical images into two separate carrier images, ensuring that the resulting container images are visually indistinguishable from the carriers. Subsequently, a dedicated recovery network extracts and reconstructs the lossy shift-rectified image and a degraded version of the original from the two carriers, using the former to calibrate and fully restore the original medical image in a lossless manner. Experimental results demonstrate the algorithm’s effectiveness in preserving privacy while enabling perfect image recovery: the container images achieve an average PSNR of 40.08 dB and MSSIM of 0.988 relative to the carriers, confirming near-perceptual-indistinguishability; the reconstructed secret images are recovered with a bit-level accuracy of 100%; and the algorithm maintains robustness under Gaussian noise, salt-and-pepper noise, and PCA compression with lossless recovery guaranteed. Moreover, the method shows excellent generalization across multiple medical imaging datasets (CT, MRI, and X-ray), highlighting its clinical applicability and innovation in secure medical image transmission systems.
The unprecedented growth of computer vision applications, such as surveillance systems and social media, raises security and visual privacy concerns, especially when data is stored on cloud servers. Image obfuscation offers a way to preserve visual privacy while maintaining an adequate level of usability; thus, it has been a topic of great interest in recent years. However, prior obfuscation schemes are either vulnerable to malicious attacks, such as model inversion to reconstruct original images from obfuscated images, or generate non-trainable obfuscated images, making them unusable for achieving reasonable accuracy. This paper proposes a novel bit-plane-based image obfuscation scheme, {\em Bit-ViP}, to preserve visual privacy for image-based recognition tasks. The Bit-ViP scheme produces secure, usable images by incorporating an innovative end-to-end obfuscation function. While doing so, the obfuscated image would contain non-invertible noise (generated by Lorenz's chaotic system and differential privacy), making it hard for an adversary to reconstruct the original image. We conduct extensive experiments on two popular activity recognition datasets, namely UCF101 and HMDB51, to validate the effectiveness of Bit-ViP. In the face of attacks on reconstruction, pixel frequency, information entropy, and pixel inter-correlation, we present a rigorous security analysis demonstrating tangible improvements over existing schemes.
V. Tanwar, Ashish Gupta, S. Madria et al.· IEEE Transactions on Emergin...· 0 citations
Generative diffusion models have revolutionized facial image synthesis, yet robust identity preservation in high resolution outputs remains a critical challenge. This issue is especially vital for security systems, biometric authentication, and privacy sensitive applications, where any drift in identity integrity can undermine trust and functionality. We introduce Diff-ID, a diffusion based framework that enforces identity consistency while delivering photorealistic quality. Central to our approach is a custom 210K image dataset synthesized from CelebA-HQ, FFHQ, and LAION-Face and captioned via a fine tuned BLIP model to bolster identity awareness during training. Diff-ID integrates ArcFace and CLIP embeddings through a dual cross attention adapter within a fine tuned Stable Diffusion UNet. To further reinforce identity fidelity, we propose a pseudo discriminator loss based on ArcFace cosine similarity with exponential timestep weighting. Experiments on held out and unseen faces show that Diff-ID does not exceed InstantID in raw ArcFace Face Similarity, but achieves substantially lower FID and the strongest FIQ based identity--realism trade off among the evaluated methods. We also present a unified DDIM based morphing pipeline that enables qualitative facial interpolation without per identity fine tuning. We further argue that identity preservation and photorealism should be evaluated jointly rather than in isolation, as high identity similarity alone does not guarantee realistic outputs. To make this trade off explicit, we report Face Image Quality (FIQ) as a complementary ratio based score that combines identity similarity and perceptual realism while keeping FS and FID as the primary metrics.
T. Rizwan, Sara Atito, Muhammad Awais et al.· 0 citations