Skip to content
Preprint

Generalized Synthetic Image Detection with Enhanced RGB-Noise Representation Learning

Jul 2026 · 0 citations
Computer Science

TL;DR

This work proposes RNSIDNet, a novel forensic framework that achieves robust detection through enhanced RGB-Noise representation learning, and designs a Hard Sample-aware Contrastive Learning (HSCL) strategy.

Abstract

The rapid advancement of large-scale generative models has accelerated the spread of highly deceptive AI-generated images, making generalized synthetic image detection a critical imperative. Existing forensic networks often struggle with cross-model generalization and realworld degradations due to their reliance on single-domain representations and conventional binary classification optimization. To overcome these limitations, we propose RNSIDNet, a novel forensic framework that achieves robust detection through enhanced RGB-Noise representation learning. Specifically, our method employs a dual-branch architecture where global RGB semantics, extracted by an attention-refined CLIP backbone, dynamically modulate highfrequency noise artifacts captured by Bayar convolutions via a Feature-wise Linear Modulation (FiLM) module. To further enhance the learned representations, we design a Hard Sample-aware Contrastive Learning (HSCL) strategy. By explicitly penalizing challenging training samples, HSCL reshapes the latent feature space to maximize the discriminative margin between pristine and synthetic domains. Extensive experiments across eight public benchmark datasets verify that our model achieves state-of-the-art performance, delivering superior generalization ability, robustness, and computational efficiency. Code and dataset will be publicly available on https://github.com/multimediaFor/RNSIDNet.

View source

Similar papers

Jul 2026

Effective Synthetic Image Detection via Noise Residual Clustering

Noise residual fingerprints are extracted by a simple yet effective pre-trained Noiseprint++ model, outperforming the state-of-the-art detectors on generalization ability, and the effectiveness of each module is validated by ablation studies.

Caihui Yan, Gang Cao, Huawei Tian et al. · 0 citations
Aug 2026

AdaMultiGAN: an adaptive multiscale decoding framework for few-shot image generation

An adaptive multi-scale decoding framework that effectively balances global context with fine-grained detail is proposed that exhibits superior robustness and generalization across diverse domains, effectively alleviating limitations of existing fusion-based approaches.

Yu Luo, Chunna Zhao, Yaqun Huang · 0 citations
#generative ai Sep 2026

MIDNet: multi-scale interaction and dynamic hard-sample mining for AI-generated image detection

With the rapid advancement of generative artificial intelligence (AI), the visual fidelity of synthesized images has increased dramatically, posing serious challenges to the verification of digital content authenticity. Existing AI-generated image detection methods often suffer from limited generalization and robustness, particularly when confronting unknown generative models or complex post-processing perturbations. To attenuate such deficiency, we propose an AI-generated image detection scheme. Leveraging a frozen contrastive language–image pre-training with Vision Transformer as visual backbone, the network extracts and stacks multi-scale intermediate features from the transformer modules to effectively capture both low-level and high-level forensic fingerprints. Based on this representation, we introduce an improved convolutional block attention module, which adopts a cascaded design by first applying channel-wise attention and then spatial attention. This design enables the network to adaptively select informative feature hierarchies while strengthening the representation of local generative artifacts. To further optimize the feature space structure, we propose a hard-sample-aware contrastive learning loss. It dynamically mines hard samples to enhance intra-class compactness and inter-class separability. In addition, we construct a mixed-source training image dataset named mixed-source AI-generated image dataset, which covers diverse generative paradigms. Large-scale testing results show that our proposed scheme ranks first on average across seven benchmark datasets, with accuracy 82.96% and the area under receiver operating characteristic curve 93.43%, demonstrating its outstanding generalization ability. Code is publicly available at https://github.com/multimediaFor/MIDNet.

Unknown authors · 0 citations
Jul 2026

Hybrid Semantic and Spectral Ensemble for Robust Synthetic Image Source Attribution

This work introduces a dual-branch ensemble framework fusing Semantic Deep Learning with Mathematical Forensic Feature Extraction, highlighting the practicality and scalability of mathematical forensics for real-world deployment.

Md. Ajwad Hossain · 0 citations
Conference Aug 2026

DKS-Net:a depthwise kernel selective network for single image dehazing

A dehazing framework named DKS-Net is proposed which fully utilizes the physics guiding features and extracting structural information in the spatial domain, and a Kernel Selective Feature Extraction Module (KSFE) is introduced to effectively captures structural patterns via large-kernel convolutions with dynamic selection capabilities and multi-scale semantic cues.

Zehao Shi, Han Wang, Xinyue Liu · 0 citations
Preprint Aug 2026

Generated Images Are Easier to Forget: A Machine Unlearning Perspective for Synthetic Image Detection

This work establishes a new paradigm for generated image detection by recasting the detection task as a problem of machine unlearning, and introduces two detection methods: data-free detection, which prunes model parameters to induce unlearning without data access, and data-driven detection, which optimizes LVMs to unlearn knowledge tied to generated images.

Jun Nie, Yonggang Zhang, Tongliang Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.