Jul 2026· 2026 International Conference on Emerging Trends in Information, Communication & Systems (ICETICS)· pp. 1-5· 0 citations· 15 references
Abstract
The pervasiveness those involving computer generated images, especially in the last decade has seen massive generational leaps, with the likes of GANs, autoencoders, and diffusion models now advanced enough to produce images so real that distinguishing between such and existing pictures becomes an uphill task. This paper serves a significant role in the improvement of modern deep learning models that can tackle different AI image generators and their performance in simulated images. The authors provide data pools with collected datasets to close previous research on gap existence such as limited cross-generator generalization, lack of focus on fine-grained relic detection, and absence of accurate alterations. The abundant dataset being discussed is a tangle of real imaginations, faked media Recognition video sequences, StyleGAN-made pictures, ProGAN/PGGAN outputs, and Stable Diffusion artificial visuals which were collected from Kaggle. Each of the dataset classes was leveled in quantity and then made rugged with different types of real-world alterations like solidity pieces, noise in low sunlit, occlusions, impression, and so on. A variety of other artifacts occurring in the output were ultra local and required more specific information for elimination. Plans were further made to create a model that would perform better with structured information based on the global context such as the frame-level status with fine-grained artifacts around and the proposed improvement focuses on the recognition of textured elements in the image on the level of individual parts. Extensive and exhaustive tests assessing the estimation features based on F1-score, precision, accuracy, recall, ROC-AUC metric as well as cross-generator estimation which involve the very behavior depreciation as a weakness revealed that the mentioned framework can be applied to new unseen generative models while preserving the proposed optimality with the diffusion-based datasets included.
The proposed Attention-Based Deep Learning Pipeline of AI-Created Image Recognition incorporates three integrated branches, including low-level statistical feature extraction, high-level semantic representation learning, and attention-based feature refinement mechanism, which support the robustness and generalization ability of the proposed model in detecting AI-generated images in a variety of generators and conditions.
Nadia Ali· Al-Noor Journal of Engineeri...· 0 citations
Recent advances in generative models like Style-GAN2 and diffusion models produce highly realistic AI-generated faces, necessitating reliable detection for identity verification and security applications. This study presents a systematic empirical comparison of Vision Transformer ViT-B/16 and EfficientNetV2-B3 via transfer learning for binary classification of real versus AI-generated faces. A combined dataset of 200,000 images was constructed, integrating real faces from CelebA and FFHQ with synthetic faces from StyleGAN2 and modern text-to-image models including Flux, DALL-E 3, and Stable Diffusion XL. Both models were initialized with ImageNet pretrained weights and optimized using a two-stage fine-tuning scheme that progressively unfreezes backbone layers. Performance was evaluated on accuracy, precision, recall, F1-score, ROC-AUC, PR-AUC, and computational efficiency metrics including latency and throughput. We further analyzed per-generator performance, model interpretability via Grad-CAM and Attention Rollout, robustness against image degradations, and cross-dataset generalization. EfficientNetV2-B3 achieved top performance with 0.9986 accuracy and F1-score, slightly outperforming ViT-B/16 (0.9979 accuracy), with both achieving near-perfect ROC-AUC. Both models maintained above 0.99 accuracy under JPEG compression and resizing. However, ViT-B/16 demonstrated superior cross-dataset generalization (0.9474 accuracy) on an external StyleGAN test set, compared to EfficientNetV2-B3 (0.8901). Since EfficientNetV2-B3 uses approximately 6.7 times fewer parameters and achieves 2.6 times higher throughput, it is recommended for in-distribution high-throughput detection, while ViT-B/16 is preferred for cross-generator scenarios.
G. Indra, K. Lhaksmana· International Conference on...· 0 citations
This project presents an explainable deep learning framework for identifying real and AI-generated images using the NASNet architecture and achieves high detection accuracy while providing interpretable visual explanations, making it suitable for digital image verification, media authentication, and cybersecurity applications.
Panduga Mounika, Dr.CH. Buchi Reddy· American Journal of AI Cyber...· 0 citations
The rapid advancement of image generation models has made it increasingly difficult for people to distinguish AI-generated images from real ones. To prevent the potential risks associated with the misuse of fake images, AI-generated image detection has gained significant attention. Existing methods neglect the inherent differences between real and fake images, thus lacking robustness and generalization ability. In this work, we innovatively investigate AI-generated image detection using bit-planes, and introduce the bit-reversed image. We propose a simple yet effective pipeline consisting of construction of bit-reversed images, gradient-based patch selection and a convolutional classifier. Besides, we provide a theoretical analysis from the mathematical perspective to demonstrate the validity of our approach. We also introduce two challenging datasets for AI-generated image detection. Extensive experiments verify the effectiveness of our approach across different settings, including cross-generator generalization, cross-dataset generalization and zero-shot performance. Without bells and whistles, our approach outperforms existing methods on over 40 benchmarks, and is nearly 100 times faster than counterparts. The code is at https://github.com/renxi-seu/RAID.
R. Cheng, Jie Gui, Hongsong Wang· arXiv.org· 1 citation
With the rapid advancement of generative artificial intelligence (AI), the visual fidelity of synthesized images has increased dramatically, posing serious challenges to the verification of digital content authenticity. Existing AI-generated image detection methods often suffer from limited generalization and robustness, particularly when confronting unknown generative models or complex post-processing perturbations. To attenuate such deficiency, we propose an AI-generated image detection scheme. Leveraging a frozen contrastive language–image pre-training with Vision Transformer as visual backbone, the network extracts and stacks multi-scale intermediate features from the transformer modules to effectively capture both low-level and high-level forensic fingerprints. Based on this representation, we introduce an improved convolutional block attention module, which adopts a cascaded design by first applying channel-wise attention and then spatial attention. This design enables the network to adaptively select informative feature hierarchies while strengthening the representation of local generative artifacts. To further optimize the feature space structure, we propose a hard-sample-aware contrastive learning loss. It dynamically mines hard samples to enhance intra-class compactness and inter-class separability. In addition, we construct a mixed-source training image dataset named mixed-source AI-generated image dataset, which covers diverse generative paradigms. Large-scale testing results show that our proposed scheme ranks first on average across seven benchmark datasets, with accuracy 82.96% and the area under receiver operating characteristic curve 93.43%, demonstrating its outstanding generalization ability. Code is publicly available at https://github.com/multimediaFor/MIDNet.
Unknown authors· Journal of Electronic Imagin...· 0 citations
Experimental results demonstrate that the proposed approach effectively identifies deepfake images with high accuracy, making it suitable for applications in digital forensics, media verification, and cybersecurity.
J. Kollu, Mortha Pavan, Putta Vardhan et al.· International Journal of Inn...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.