Multimodal representation learning is critical for a wide range of applications, such as multimodal sentiment analysis. Current multimodal representation learning methods mainly focus on the multimodal alignment or fusion strategies, such that the complementary and consistent information among heterogeneous modalities can be fully explored. However, they mistakenly treat the uncertainty noise within each modality as the complementary information, failing to simultaneously leverage both consistent and complementary information while eliminating the aleatoric uncertainty within each modality. To address this issue, we propose a plug-and-play feature causality decomposition method for multimodal representation learning from causality perspective, which can be integrated into existing models with no affects on the original model structures. Specifically, to deal with the heterogeneity and consistency, according to whether it can be aligned with other modalities, the unimodal feature is first disentangled into two parts: modality-invariant (the synergistic information shared by all heterogeneous modalities) and modality-specific part. To deal with complementarity and uncertainty, the modality-specific part is further decomposed into unique and redundant features, where the redundant feature is removed and the unique feature is reserved based on the backdoor-adjustment. The effectiveness of noise removal is supported by causality theory. Finally, the task-related information, including both synergistic and unique components, is further fed to the original fusion module to obtain the final multimodal representations. Extensive experiments show the effectiveness of our proposed strategies.
Ye Liu, Zihan Ji, Hongmin Cai· Neural Information Processin...· 3 citations
Spatial transcriptomics (ST) enables the simultaneous measurement of high-throughput gene expression and spatial structural information, offering a powerful means to decipher tissue heterogeneity. However, current spatial domain identification methods struggle to accurately distinguish continuous biological boundaries, such as smooth tissue transitions or invasive tumor margins. To overcome this challenge, we propose ENGGT, an edge-node guided graph transformer framework that identifies spatial domains with accurate biological boundaries by modeling interactions between node and edge representations. ENGGT employs a dual-branch architecture to capture spatial patterns. Specifically, an edge transformer branch encodes edge features to characterize the spatial and expression relationships of spots pairs within a local tissue microenvironment. It incorporates a topology-aware edge masking strategy to prune unreliable connections and enhance boundary sensitivity. A node transformer branch then integrates the learned edge representations as local attention biases into a global self-attention module, promoting intra-region consistency while limiting error propagation across biological boundaries. Evaluated on ST datasets from multiple measurement platforms, ENGGT consistently outperforms seven state-of-the-art spatial domain identification methods across all evaluation metrics. Moreover, the learned edge-bias matrix offers traceable biological interpretability and enables biological boundary localization. ENGGT provides a robust and interpretable tool for spatial transcriptomics analysis by identifying spatially coherent tissue structures and delineating precise boundaries.
Jiazhou Chen, Ziru Xiao, Junyu Li et al.· Proceedings of the 32nd ACM...· 0 citations
Semi-supervised learning (SSL) aims to effectively utilize a small amount of labeled data together with a large volume of unlabeled data to improve learning performance. Among various SSL strategies, label propagation has been widely adopted due to its ability to diffuse label information across data points via graph structures. However, most existing label propagation-based SSL methods struggle with high-dimension-low-sample-size (HDLSS) data, as they rely on pairwise similarity measures that fail to capture the complex relationships among samples in such settings. To overcome this limitation, we propose a novel deep SSL framework that enhances label propagation using tensor-based similarity, enabling the modeling of high-order relationships among multiple samples. Specifically, we first pretrain a feature extraction network using the labeled data to obtain initial feature representations. Subsequently, tensor label propagation and fine-tuning of the feature extraction network are conducted iteratively. In the tensor label propagation module, pseudo-labels for the unlabeled samples are estimated more accurately by leveraging high-order similarity. These pseudo-labels, along with the original labeled data, are then used to fine-tune the pretrained network, resulting in more robust and discriminative feature representations across all samples. By embedding high-order structural information into the SSL pipeline, our method significantly enhances the prediction performance from limited labeled data in the HDLSS setting. Extensive experiments on multiple HDLSS datasets demonstrate the superiority of our approach compared to recent baselines.
Hongmin Cai, Jiali Sun, Fei Qi et al.· IEEE Transactions on Neural...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.