Spatial-Frequency Fusion for Robust Detection of AI- Generated Images across Unseen Generators
The rapid development of generative artificial intelligence has made it increasingly difficult to distinguish AI-generated images from authentic photographs. Generative adversarial networks, diffusion models, and text-to-image systems can produce highly realistic visual content, creating challenges for digital forensics, misinformation control, and content authentication. Existing detectors frequently depend on generator-specific artifacts and may lose effectiveness when images originate from unseen generators or undergo resizing, compression, or blur. This study proposes a spatial–frequency feature-fusion framework for binary detection of real and AI-generated images. A convolutional neural network extracts spatial and structural representations from RGB images, while a second branch analyzes frequency-domain characteristics obtained through a two-dimensional Fourier transform. The resulting representations are fused before classification. Data augmentation is incorporated to improve robustness against common image transformations. The framework is designed for evaluation using the GenImage benchmark, which contains more than one million real/fake image pairs and explicitly supports cross-generator and degraded-image evaluation [10]. Performance is assessed using accuracy, precision, recall, F1-score, and AUROC. The experimental design emphasizes generalization rather than relying only on in-distribution accuracy.