Aug 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· 0 citations· 10 references
TL;DR
Experiments show that LAUA substantially mitigates performance degradation under modality missingness across retrieval and regression tasks, attaining up to 20% relative improvement in MRR for retrieval and up to 24.3% relative improvement in MSE for regression.
Abstract
Multimodal federated learning enables multiple clients to collaboratively train models from distributed multimodal data while preserving data privacy. In realistic federated settings, multimodal samples are often missing or unpaired, and cross-modal heterogeneity across clients can hinder stable optimization. Many existing federated multimodal methods attempt to mitigate modality missingness by generating synthetic paired data through data augmentation or generative models. However, they typically rely on paired supervision or treat client updates uniformly, making them brittle under modality missingness and client-level variability. To address these challenges, we propose a federated multimodal learning framework (LAUA) for learning from a mixture of unimodal and multimodal clients. On clients, LAUA aligns representations in a shared variational latent space, where KL regularization yields a principled and lightweight confidence signal for estimating uncertainty. Unimodal clients learn transferable representations via self-supervised objectives, while multimodal clients additionally leverage task supervision and incorporate an internal distillation component to enhance cross-modal consistency and stabilize local optimization. On the server, LAUA performs uncertainty-weighted aggregation that adaptively down-weights unreliable client updates. Experiments on various datasets show that LAUA substantially mitigates performance degradation under modality missingness across retrieval and regression tasks, attaining up to 20% relative improvement in MRR for retrieval and up to 24.3% relative improvement in MSE for regression.
FedTaste is proposed, a parameter-efficient framework for topology-aware structural transfer in Multimodal Federated Learning with missing modalities that avoids explicit modality imputation while preserving shared semantic structure across clients.
Haocheng Liang, Jie Zhang, H. Ochiai· arXiv.org· 0 citations
Flux is proposed, a multimodal federated learning framework built around two complementary components, modality-aware confidence tempering and gradient-decoupled private adaptation, that enables sample-specific, client-local confidence adaptation without allowing confidence-dependent gradients to perturb shared representation learning.
Adiba Orzikulova, Jaehyun Kwak, Jaemin Shin et al.· 0 citations
The metrics of individual modality contribution (IMC) and multimodal synergistic gain (MSG) are introduced to quantify sample-level and semantic-level utility, so as to guide semantic denoising selection and robust conditional balancing strategies, effectively mitigating noise interference.
Yan Zhang, Xiaoye Miao, Yanming Yu et al.· Proceedings of the 32nd ACM...· 0 citations
Experiments show that FedAMB improves multimodal accuracy and missing-modality robustness and prevents the propagation of fusion-biased teachers while directly improving unimodal representations.
Seung-Hwa Han, Juyeob Lee, Sang-Min Lee et al.· 0 citations
ProMoE-FL, a Prototype-conditioned Mixture-of-Experts framework for robust missing-modality feature synthesis in multimodal federated learning with missing modality, builds a global client-aware prototype bank that captures clinically meaningful modality priors across institutions.
Aavash Chhetri, Bibek Niroula, Eduard Vazquez et al.· 0 citations
Federated Continual Multimodal Learning (FedCMM), a framework that embeds continual-learning safeguards into the federated optimization loop at three complementary levels, is proposed, confirming that holistic, modality-aware optimization enables robust evolutive adaptation across heterogeneous networked AI deployments.