Skip to content
Open access

Learning Generalizable Visual Representations for AI-Generated Image Detection

Jul 2026 · Al-Noor Journal of Engineering Management and Computer Science · 0 citations · 3 references

Abstract

Detectors of AI generated images often perform extremely well when learning from generators they have never encountered information from, and fail when learning from generators they have not learned from. Detectors of AI generated images typically have an accuracy rate of almost 100% when they are trained on generators they have not seen information from, but perform poorly when they are trained on generators they have not seen information from, or when they are trained on a visual subject matter they have not seen information from. This paper poses the question of which visual representations actually convey transferable evidence of synthetic image formation. Using a controlled corpus of 217,100 images spanning twelve generators from five architectural families (style-based GANs, latent diffusion, pixel-space diffusion, autoregressive transformers and closed commercial systems) and five semantic domains, in which real and generated images are matched on prompt, resolution, aspect ratio, format and compression and de-duplicated by perceptual hashing and embedding similarity, we benchmark nine representation families under a common parameter budget. We then propose three-stream detector, GVR-Det, which incorporates a semantic-structural, a local-forensic, and a spectral encoder, all of which are connected by a gate, and which is trained by class-conditional generator and domain-invariance, cross-generator contrastive alignment and transformation consistency. In addition to accuracy we quantify generator leakage, domain leakage, class-conditional maximum mean discrepancy and centred kernel alignment, under leave-one-generator-out, leave-one-family-out, leave-one-domain-out and joint protocols. The best transferring representation is not the best distributing representation: there is significant domain information in the global semantic features, while the residual and spectral features are more stable across the generator families, but more vulnerable to compression. The gated, invariance-regularised model is the best joint out-of-distribution behaviour and its generator leakage only drops from 0.914 to 0.402; it degrades the least when JPEG re-encoded, rescaled and blurred.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.