Radial Residual Frequency: A Semantically Aligned Benchmark and Spectral Detector for AI-Generated Images
Abstract
We study the detection of AI-generated images and contribute toward detectors that are accurate, trustworthy, and honestly evaluated. We first describe a data-generation pipeline that captions real photographs with a vision–language model and regenerates them with modern text-to-image systems, producing semantically aligned real/synthetic pairs that isolate generative artifacts from image content. We then build a lightweight, CPU-deployable spectral detector that fuses an RGB backbone with a radial residual-frequency branch, and show through a controlled ablation that the frequency cue mainly contributes calibration and false-positive control rather than raw ranking. To keep a single saturated score from overstating readiness, we package the detector with a multi-axis evaluation suite and a worst-group summary metric. Finally, we extend the same recipe to a dedicated face-deepfake detector that adds a neighboring-pixel-relationship branch and generalizes well to generators unseen in training. We release the models, the data protocols, and all per-sample scores as a reproducible reference point for synthetic-image forensics.