Neural Codec Resonance: Investigating Zero-Shot Generative Music and Speech Forensics via Multi-Resolution Spectral and Phase Invariances
Abstract
odels—spanning neural vocoders, latent diffusion architectures,and discrete neural codecs—have enabled the synthesis ofhighly realistic musical pieces and vocal performances. Whiledata-driven classifiers frequently overfit to superficial datasetartifacts, physically grounded methods aim to detect the invariantmathematical footprints imposed by the underlying synthesisarchitectures. In this work, we investigate Neural Codec Resonance(NCR), an analytical framework examining three physical andstructural properties of synthetic audio: (1) ultrasonic codebookbandwidth roll-offs caused by Residual Vector Quantization (RVQ)decimation filters, (2) stereo phase dispersion governed by physicalwave propagation acoustics (the Haas effect) versus pseudo-stereodiffusion collapse, and (3) reconstruction residuals under multi-resolution short-time Fourier transform (STFT) autoencoderprojections.Through empirical evaluations across two distinct modalities,we delineate the practical capabilities and boundary conditions ofthese representations: In polyphonic music evaluation (N = 600tracks across five transmission channels: Clean WAV, MP3320 kbps, MP3 128 kbps, AAC 256 kbps, and Opus 16 kbps),stereo phase coherence coupled with ultrasonic boundary trackingachieves robust discrimination between neural music models (Suno,Udio, MusicGen) and acoustic orchestral/jazz studio recordings.Conversely, when evaluating single-channel speech deepfakesin unconstrained real-world settings (N = 60 authentic andsynthetic speech files from open benchmarks), static harmoniccomb heuristics experience substantial degradation (AUROC of0.2844, detection recall of 30.00%, and a false accusation rateof 10.00%), demonstrating that dynamic vocoder pitch tracking,acoustic background noise, and mastered loudness floors confoundsimple spectral peak models. We discuss the implications of thesefindings for automated catalog moderation and voice verificationsystems.Index Terms—Audio Forensics, Generative Music, Neural AudioCodecs, Residual Vector Quantization, Haas Effect, Model ContextProtocol.