This work introduces TD-V2A, which leverages temporal differences (TD) as the key representation that distinguishes V2A from I2A, enriching visual conditioning with minimal architectural modification and significantly improves end-to-end V2A generation quality, even outperforming dedicated V2A representations such as contrastive audio-visual pretraining.
Abstract
Video-to-audio (V2A) generation extends image-to-audio generation (I2A) by introducing consecutive frames that provide essential temporal cues for audio synthesis. However, existing conditional diffusion-based V2A methods typically enhance visual conditioning with additional audio-visual supervision, acoustic structure prediction, or reasoning from large multimodal models, requiring extra networks or strong inductive biases. Inspired by recent advances in visual representation learning, we introduce TD-V2A, which leverages temporal differences (TD) as the key representation that distinguishes V2A from I2A, enriching visual conditioning with minimal architectural modification. We first investigate TD at both the frame and feature levels to identify the most effective representation level at which TD complements visual representations. Based on these findings, we develop a hierarchically continual learning strategy and an annealed temporal differences guidance method to progressively learn and exploit TD information during diffusion training and sampling process, respectively. Extensive experiments on benchmark datasets demonstrate that effectively exploiting TD through our proposed framework significantly improves end-to-end V2A generation quality, even outperforming dedicated V2A representations such as contrastive audio-visual pretraining.
This framework produces both a transformation embedding and a processed-audio embedding, and it finds that the two play complementary roles: distance-based tasks favor the former, while probe-based tasks favor the latter.
Sungho Lee, Marco A. Mart'inez-Ram'irez, Junghyun Koo et al.· 0 citations
OmniVAE is presented, a jointly trained audio-video VAE that learns fine-grained semantic alignment between audio and video latent representations that translates into higher generation quality and more accurate cross-modal synchronization in downstream text-to-audio-video generation.
Jun Zhan, Chenchen Yang, Yitian Gong et al.· arXiv.org· 0 citations
Detecting audio-visual DeepFake (AVDeepFake) is becoming increasingly important as synthetic media tools become widely accessible and spread across consumer devices. In this study, we present an Detecting audio-visual DeepFakes (AV-DeepFakes) has become increasingly critical with the rapid proliferation of accessible synthetic media generation tools across consumer platforms. In this work, we propose a highperformance, deployment-efficient AV-DeepFake detection framework tailored for real-world consumer devices. The proposed model integrates a 3D convolutional visual encoder with a 2D convolutional audio encoder to learn synchronized multimodal representations, effectively capturing spatial, spectral, and prosodic inconsistencies inherent in manipulated content. To detect temporal forgeries, we introduce a bidirectional complementary boundary module that precisely localizes manipulation onsets and offsets. A cross-modal attention fusion mechanism aggregates modality-specific cues, while an uncertainty-aware gating strategy suppresses unreliable signals to improve robustness. Furthermore, a cross-modal discrepancy minimization loss encourages alignment for genuine samples while maximizing divergence for forged content, strengthening multimodal consistency learning. Extensive evaluations on FaceForensics++ and LAV-DF demonstrate the effectiveness of the proposed approach, achieving 97.1% AUC for clip-level detection and an 81.2% F1-score for temporal boundary localization, while reducing inference time by 5× compared with transformer-based methods.
Nasir Saleem, Adeel Hussain, Sami Bourouis et al.· International Journal of Int...· 0 citations
This work presents a synchronization-aware acceleration framework for efficient audio-visual generation by explicitly accounting for cross-modal dependence during acceleration, and improves inference efficiency while keeping video quality, audio quality, and audio-video synchronization.
Sheng-Chuan Gao, Teng Hu, Bohao Feng et al.· 1 citation
DSF-Net is a novel neural network that introduces a dual-strategy fusion approach for the AVSELD task, designed for robust and computationally efficient multi-modal comprehension.
The core of SSVAL is Visual Anchor Prompt Injection (VAPI), which introduces prompts that absorb rich knowledge from external VFMs during training, enabling them to serve as stable visual anchors that mitigate representation deviation during inference.
Qian-Long Yang, Bowen Ye, Xianda Guo et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.