The proposed GlobalForge improves average BAcc on 8 in-the-wild benchmark groups by $\mathbf{5.89\%}$ over the previous state-of-the-art, and is clearly ahead of representative baselines on RealDeg-Bench under both single and compound degradations.
Abstract
AI-generated image (AIGI) detectors achieve strong accuracy on clean benchmarks, but their performance drops sharply after images are propagated through real-world channels. We trace this fragility to what these detectors actually learn: they overfit to local artifacts left by generators in small spatial neighborhoods, which are easily destroyed by common propagation degradations such as JPEG compression and blur. Instead, we shift the discriminative cue from fragile local artifacts to more robust global structure. Building on this, we propose GlobalForge, a framework with two complementary modules. The Local Information Bottleneck (LIB) suppresses local components to block shortcut learning, while the Global Structural Reasoning (GSR) module forces every token to gather evidence from distant regions. Both modules are trained jointly under a contrastive structural loss based on degradation that keeps the resulting features stable under degradation. To support fine-grained robustness evaluation, we further introduce RealDeg-Bench, covering 7 common degradation operators and multi-step compound chains. GlobalForge improves average BAcc on 8 in-the-wild benchmark groups by $\mathbf{5.89\%}$ over the previous state-of-the-art, and is clearly ahead of representative baselines on RealDeg-Bench under both single and compound degradations. Code is available at https://anonymous.4open.science/r/GlobalForge-BE0F/.
AI-generated content (AIGC) has become increasingly difficult to distinguish from real images, creating new challenges for media authentication. Existing detectors often rely on either convolutional networks, which focus on local patterns but lack global reasoning, or Transformers, which capture long-range context but suffer from high computational cost. Recent state space models such as Mamba provide linear-time processing, yet their causal structure leads to long-range dependency decay, making them less effective for detecting forgery clues that appear across distant image regions. In this work, we propose Multi-scale Linear Local Attention (MLLA), a unified framework for AIGC detection that combines local artifact modeling with efficient global context reasoning. Our design integrates an artifact-aware tokenization (AAT) with a Linear Local Attention (LLA) block that merges depthwise convolutions, linear attention, and rotary positional embedding to overcome the limitations of both Transformers and causal state space models. By stacking LLA blocks in a multi-scale encoder, the network learns fine-grained features in shallow layers and broader semantic clues in deeper layers. Experiments on a wide range of GAN and diffusion datasets show that MLLA achieves state-of-the-art performance and strong generalization to unseen generators. The results confirm that combining local priors with efficient non-causal global modeling is a simple yet powerful direction for AIGC detection.
Xiao-Long Liu, Mengyao Xiao, Haorui Wu et al.· IEEE Signal Processing Lette...· 0 citations
This work investigates what cues are exploited by foundation-model-based detectors to distinguish real images from diffusion-generated ones and suggests that foundation-model-based detectors succeed by capturing non-semantic low-to-mid frequency distributional discrepancies between real and diffusion-generated images.
SSMB is presented, a deblur-free, self-supervised keypoint detector for motion-blurred images that requires neither handcrafted detectors nor external pseudo-labels, and introduces the Local Discriminability Enhancement (LDE) module, which restores fine-grained local discriminability after global feature mixing.
Zhenjun Zhao, F. Bellavia, Wen-Ting Wang et al.· 0 citations
Qualitative analysis suggests that PatchHead reduces class-conditional domain discrepancy, and redirects the representation from content-dominated saliency toward spatially distributed authenticity evidence, which provides a representation-level account of why spatial patch aggregation transfers more reliably across generators and datasets than a single CLS-based global representation.
Shengbo Qi, Hongyi Fang, Benjia Zhou et al.· 0 citations
Generalization remains a critical bottleneck in AI-generated image detection. Because many modern generators are proprietary or adversarially modified, existing detectors overfit to the low-level textural patterns of accessible training data, resulting in severe failures on unseen domains. Conventional regularization techniques (e.g., $L_1$/$L_2$ norms, Dropout) apply indiscriminate parametric constraints and fail to provide the domain-invariant structure necessary for cross-generator robustness. To address this, we propose Feature-Augmented Implicit Regularization (FAIR). FAIR introduces an orthogonal, macro-structural prior, specifically, Scene Composition Structure (SCS), during training to geometrically constrain the model's optimization trajectory. By augmenting the primary feature space with domain-invariant SCS features, FAIR explicitly penalizes texture-biased shortcut learning. Crucially, this structural prior is entirely discarded at inference, yielding a smoothed, generalized decision boundary with zero architectural or computational overhead. Extensive evaluations across five massive benchmarks demonstrate that integrating FAIR into state-of-the-art detectors significantly improves cross-generator generalization, boosting accuracy by up to 8.04% and establishing new state-of-the-art robustness in zero-shot transfer scenarios.
Md Redwanul Haque, M. Murshed, Manoranjan Paul et al.· arXiv.org· 0 citations
FIT employs a lightweight Degradation Encoder to predict a global degradation vector and a spatial degradation map from local degradation severity, which jointly condition the patch embedding and unembedding through adaptive deformation, and introduces a task-token dropout strategy that regularizes task conditioning during training.
Zihao He, Yunfeng Wu, Xinchao Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.