Results establish that dominance-aware control and missing-safe fusion are effective principles for robust multimodal recognition under structural modality imbalance.
Abstract
Modality imbalance undermines multimodal pattern recognition when auxiliary inputs are missing, noisy, or weakly aligned, causing fusion models to collapse toward quasi-unimodal shortcuts and produce unreliable confidence. We propose DOMINO, a dominance-aware framework that mitigates this failure through three mechanisms: (1) a closed-loop dominance controller that regulates auxiliary objectives during training, (2) a missing-safe residual fusion module that enables stable fallback under degraded modalities, and (3) reliability-aware inference via temperature scaling and weighted expert aggregation. Experiments on Weibo and MediaEval2015 demonstrate state-of-the-art performance, achieving 0.9785 macro-F1 on Weibo and 0.9559 macro-F1 on MediaEval2015, while substantially reducing expected calibration error compared with that of strong baselines. These results establish that dominance-aware control and missing-safe fusion are effective principles for robust multimodal recognition under structural modality imbalance.
GAUGE is proposed, a lightweight counterfactual gating framework for incomplete multimodal classification that outperforms strong baselines across diverse incomplete-input settings and is established as a principled and scalable framework for fine-grained evidence control under modality incompleteness.
Yun Shi, Enshui Yu, Kai-Rui Guo et al.· 0 citations
Modality Subspace Activation (MSA) is proposed, a training-free inference-time framework that uses Singular Value Decomposition (SVD) to estimate modal activation strengths and dynamically balances modal projections in the last hidden state, effectively restoring CMS across benchmarks.
Hongbo Jiang, Jie Li, Yunhang Shen et al.· 0 citations
HalluPrism, a behavioral diagnostic that re-runs an answer after visual degradation, blank-image replacement, and grounding or relation checks is proposed, a behavioral diagnostic that separates failure diagnosis from abstention scoring.
Aman Prakash, Sourish Dasgupta, Tanmoy Chakraborty· 0 citations
Multimodal intent recognition combines linguistic, acoustic, and visual evidence, but individual modalities may be noisy, missing, semantically conflicting, or disproportionately dominant. Existing methods typically infer modality importance implicitly and either reweight or suppress unreliable inputs, without determin...
Suraj Kumar, Mohnish Raj, Soumi Chattopadhayay et al.· 0 citations
Fusing multiple modalities is expected to improve model performance. However, on the MultiHuSE dataset, early, late, and symmetric attention fusion often fail to outperform the best unimodal baseline (text). Pathway isolation of a symmetric attention fusion model reveals that the text-pathway accuracy drops from 74.9%...
Mary Ogbuka Kenneth, Foaad Khosmood, Abbas Edalat· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.