Combining CRNN Modeling to Dynamically Align the Relationship Between Music Rhythm and Plot Twists in Different Film Genres
Abstract
This paper addresses the challenge of quantitatively modeling the dynamic alignment between musical rhythm and plot turning points in films by proposing V2A-AlignNet, a genre-aware cross-modal deep learning framework. As intelligent multimedia perception increasingly relies on advanced signal processing and multimodal information fusion techniques that are conceptually relevant to electromagnetic sensing and communication systems, accurate temporal alignment has become an important research topic. The proposed model adopts a dual-stream architecture integrating VideoMAE for long-range spatiotemporal video representation and a CRNN for extracting both local and global rhythmic characteristics from Mel spectrograms. A cross-modal attention module constructs a shared semantic space to generate alignment saliency sequences and similarity matrices, while a genre-conditioning mechanism enables adaptive modeling for different film categories. Experiments conducted on 120 films spanning six genres (800 clips) demonstrate that V2A-AlignNet achieves superior performance in turning-point detection, alignment pattern classification, and genre recognition compared with representative baseline methods. Ablation studies further verify the effectiveness of each component, and visualization results reveal distinctive genre-specific alignment behaviors. The proposed framework provides a computational basis for audio-visual temporal relationship analysis and offers valuable insights for multimodal signal interpretation and intelligent information processing in advanced electromagnetic sensing and communication-related applications.