MST-former: multiparty shared representation enabled transformer for secure medical image analysis
Abstract
Privacy-preserving collaborative medical image analysis across multiple healthcare institutions remains challenging due to data heterogeneity, patient privacy constraints, and incomplete cross-institutional record alignment. To address these limitations, this paper proposes MST-Former, a multiparty shared transformer framework that enables secure embedding fusion without sharing raw medical images. The framework combines a deep reinforced autoencoder (DRA) for local representation learning with adaptive positional encoding, dynamic attention masking, and transformer-based embedding aggregation to facilitate robust cross-institutional knowledge integration. In addition, a contrastive learning strategy is employed to enhance representation consistency across heterogeneous imaging sources. Extensive experiments conducted on multiple benchmark medical imaging datasets demonstrate that MST-Former consistently outperforms existing methods in segmentation and prediction performance while preserving data confidentiality. Across multiple medical image segmentation datasets, MST-Former consistently achieved strong performance, with mDice scores ranging from 92.73 to 94.95% and mIoU scores ranging from 87.31 to 92.65%, demonstrating its effectiveness and generalization capability. These results highlight the effectiveness of MST-Former as a scalable and privacy-preserving solution for collaborative medical image analysis in distributed healthcare environments.