Balanced Multimodal Federated Learning: An Efficient and Noise-Resilient Approach
Multimodal federated learning (MFL) enables the collaborative training of models across multiple modalities to achieve high predictive accuracy while preserving data privacy. However, modality imbalance remains a critical bottleneck, preventing models from attaining the theoretical performance ceiling achievable through joint training. Existing methods typically rely on the assumption of clean multimodal data, thus failing to distinguish informative hard samples from detrimental noise (e.g., modal-specific and cross-modal noise). Moreover, their prohibitive computational costs render them impractical for real-world deployment. In this paper, we propose MFedFAIR, an efficient and noise-resilient multimodal federated learning framework. We introduce the metrics of individual modality contribution (IMC) and multimodal synergistic gain (MSG) to quantify sample-level and semantic-level utility, so as to guide semantic denoising selection and robust conditional balancing strategies, effectively mitigating noise interference. MFedFAIR maintains computational efficiency by deriving metrics solely via forward propagation. Furthermore, it optimizes resource utilization by prioritizing informative, semantically aligned samples for local training and ensuring robust aggregation via quality-aware collaboration. Extensive experiments on four benchmark datasets demonstrate that MFedFAIR significantly outperforms state-of-the-art baselines in both effectiveness and efficiency within realistic noisy MFL environments.