Skip to content

Robust Audio-Visual Question Answering with Missing Modality in Training and Testing.

Aug 2026 · IEEE Transactions on Pattern Analysis and Machine Intelligence · Vol PP, pp. 1-18 · 0 citations
Medicine

TL;DR

An AVQA-specific two-stage framework that adapts established cross-modal reconstruction and dependency-modeling principles to the supervision constraints of TM-AVQA is developed, a setting in which modality-complete, audio-missing, and visual-missing samples may occur during both training and testing.

Abstract

Audio-Visual Question Answering (AVQA) requires reasoning over temporally evolving audio and visual signals to answer natural-language questions about dynamic scenes. Most existing methods assume that both modalities are available during training and testing. In practice, however, an audio or visual stream may be unavailable because of passive signal loss, such as hardware or transmission failures, or intentional removal, such as withholding visual information for privacy. We formulate Training-time Modality-Missing AVQA (TM-AVQA), a setting in which modality-complete, audio-missing, and visual-missing samples may occur during both training and testing. To address this new setting, we develop an AVQA-specific two-stage framework that adapts established cross-modal reconstruction and dependency-modeling principles to the supervision constraints of TM-AVQA. In Stage-I, a reconstruction network infers task-oriented feature-level representations of the missing modality from the available modality and question text. Its Dense Temporal-scale Reconstruction (DTR) module aggregates complementary information across multiple temporal granularities, while its Multimodal Dependency Modeling (MDM) module captures dependencies among the audio, visual, and question modalities. To supervise reconstruction when the target-modality features are unavailable, we adapt two complementary learning objectives to TM-AVQA: Cross-Modal Relation-based Contrastive Learning (CMR-CL), which exploits within-video and cross-video multimodal relations, and Cross-Sample Relation-based Pseudo-label Learning (CSR-PL), which constructs feature-level pseudo targets from relevant modality-complete samples. In Stage-II, the frozen reconstruction network is combined with existing AVQA backbones for answer prediction. We construct controlled TM-AVQA variants of MUSIC-AVQA, MUSIC-AVQA-R, and AVQA datasets by deleting one modality from selected samples. Experiments across multiple AVQA backbones, missing-modality conditions, and missing rates show consistent and robust improvements over the evaluated AVQA and missing-modality baselines. Additional ablations and analyses examine the contribution of the proposed task-specific designs, reconstruction quality, efficiency, and transfer to audio-visual recognition task.

View source

Similar papers

Preprint Aug 2026

SCoPE: Training-Free Audio-Visual Event Perception via Sparse Cross-Modal Prior Exchange

SCoPE is introduced, a training-free framework in which all queried labels compete for shared evidence and each modality guides event selection in the other, and derives an exact condition for when this competition removes an FCA in a two-label fit.

J. Jeong, Junho Yoon, Hyunju Kim et al. · 0 citations
#machine learning Preprint Sep 2026

If You Hear It, Help Find It: Orthogonal Knowledge Distillation for Open-Vocabulary Audio-Visual Event Localization

Open-vocabulary audio-visual event localization (OV-AVEL) grounds a text-queried event in time from video, audio, and language. The supervision sources available to this task can differ in temporal-boundary reliability: on OV-AVEBench, our configured visual teacher gives more reliable boundary cues than the configured...

Yi Xu, Cheng Chen, Wen-Zhuo Lei · 0 citations
Aug 2026

Cross-modal alignment enhancement for lightweight large vision language models

A Low-Complexity Cross-Modal Alignment via Projection (LCAP) network is proposed, which introduces Projective Token Compression (PTC), which leverages Mish activation and adaptive average pooling to reduce feature redundancy while enhancing discriminative information, and Positional Spatial Enhancement (PSE), which exp...

Yu-Chen Sha, Lingli Wan, Ge Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Tracing Audio Grounding and Answer Selection in Audio LLMs

Audio Large Language Models (Audio LLMs) have advanced in audio understanding, yet they can still predict the answer by reasoning from textual cues or linguistic priors rather than the provided audio. A common remedy is to train models on data whose answers cannot be inferred from text alone. This approach can improve...

Hyebin Cho, Suho Yoo, Jihoo Jung et al. · 0 citations
Preprint Aug 2026

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models

ST-Omni-R1 is proposed, which integrates FOA-derived semantic and trajectory representations with panoramic visual context and is trained through progressive curriculum learning and reasoning-tree reinforcement learning, and results on three public spatial-audio benchmarks indicate that its learned spatial and motion r...

Zhi Zeng, Cheng Zhang, Ze-Sheng Yang et al. · 2 citations
Preprint Sep 2026

Video-HolmesV2: Can MLLMs Reason with Spatio-Temporal Audio-Visual Evidence in Long Videos?

Multimodal Large Language Models have demonstrated impressive video understanding, yet their ability to reason over long-form narratives is often masked by visual-centric evaluations and inefficient context processing. Existing benchmarks over-rely on visual heuristics while marginalizing auditory cues, effectively red...

Zhao-Yang Wei, Zipeng Wang, Yu-She Cao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.