Fed-MoSeC: A Federated Learning Framework for Cross-Modal Semantic Communication Systems in Mobile Networks
Semantic communication systems in mobile networks necessitate real-time collaborative updates of semantic encoders as a result of node mobility that induces semantic extraction drift. However, diversity within modal transmission content and heterogeneous encoder architectures, along with challenges such as imbalanced training requirements and pseudo-label noise, limit the effectiveness of general collaborative update approaches. In this paper, we propose Fed-MoSeC, a novel federated learning framework for updating cross-modal semantic encoders. Our framework trains only a newly designed graph neural network-based adapter while freezing heterogeneous encoders for various modalities, converting heterogeneous cross-modal updates into a homogeneous aggregation task, and significantly reducing communication overhead. By combining confidence-based filtering with similarity-matrix distillation, the novel integrated Pseudo-label Noise Counteracting Component (PNCC) is designed to be robust to noisy data. The Training Optimization Component (TOC) based on a bi-level Cournot–Stackelberg game theoretical algorithm achieves near-optimal Subgame-Perfect Nash Equilibrium (SPNE) to incentivize across nodes and maximize update nodes’ utilities with different training levels. Experimental results highlight the advantages of Fed-MoSeC over existing potential application algorithms, reducing communication by 95–97% and improving RSUM by 18% at 70% pseudo-label noise.