A multimodal dual-contrastive learning framework for peptide property prediction is proposed, which improves both the structural encoder and the contrastive learning strategy to enhance the quality of joint sequence-structure representations.
Abstract
Peptides play important roles in biological processes and biomedical applications, and their hemolytic (Hemo) and nonfouling (NF) properties directly affect their safety and translational potential. Therefore, accurate predictive models are essential for the rational design of functional peptides. Although existing multimodal peptide property prediction methods can jointly exploit sequence and structural information, their structural encoders still rely primarily on local graph convolution and their contrastive objectives are largely focused on cross-modal alignment. Consequently, they remain limited in modeling long-range structural dependencies and in enhancing intramodal discriminability. To address these limitations, we propose a multimodal dual-contrastive learning framework for peptide property prediction, which improves both the structural encoder and the contrastive learning strategy to enhance the quality of joint sequence-structure representations. Specifically, ProtBERT is adopted as the sequence encoder, and a hierarchical GNN-Transformer structural encoder is constructed to capture local topological patterns and long-range structural dependencies. In addition, a parallel graph spatial channel attention module is introduced to enhance task-relevant structural features. Within a shared embedding space, we further design an interintra hybrid supervised contrastive learning strategy to jointly optimize sequence-structure alignment and intramodal class discriminability. Experimental results show that the proposed method achieves overall performance superior to baseline models on both hemolysis and NF prediction tasks, providing an effective framework for multimodal representation learning in peptide-property prediction.
ABSTRACT Peptides combine the favorable pharmacokinetics of small molecules with the high specificity of biologics, making them promising therapeutics. Incorporating non‐canonical amino acids (ncAAs) further enhances drug‐like properties, yet modeling remains challenging due to chemically modified residues and combinatorial sequence diversity. Here, we introduce SinCAA, a similarity‐enhanced pretraining framework specifically designed to encode ncAAs. The framework is built on the principle that amino acids with similar 3D conformations induce minimal perturbations to peptide properties. It jointly optimizes two complementary self‐supervised tasks: contrastive learning guided by a conformational similarity metric to capture functional relationships among ncAAs, and masked node reconstruction to encode the unique chemical identity of each ncAA. Built on a graph transformer backbone, this dual “relationship–identity” supervision enables SinCAA to learn robust atomic representations that generalize from individual ncAA building blocks to full‐length peptides. SinCAA exhibits strong zero‐shot performance in peptide property prediction and consistently outperforms state‐of‐the‐art pretrained models across diverse benchmarks. This framework provides an efficient and interpretable approach for in silico prediction and ranking of ncAA‐containing peptides, accelerating candidate screening in therapeutic peptide discovery.
Chen-Cheng Xu, Lesong Wei, Jian-Min Wang et al.· Advancement of science· 0 citations
This work proposes GraphTransDTI, a synergistic hybrid framework that integrates a Graph Transformer to represent drug graph structures, a CNN-BiLSTM network to encode protein sequence context, and a Cross-Attention mechanism to model cross-domain interactions.
Vang V. Le, Mai Thi Anh Nhu, Pham Truong Viet Thong· PLoS ONE· 0 citations
TPT provides a compact and computationally efficient encoder with competitive performance for microbial peptide analysis and suggests that effective smORF representation may benefit from pretraining objectives and inductive biases tailored to short, rapidly evolving sequences, rather than from model scale alone.
This paper proposes TextDTI, a multimodal framework that simultaneously exploits sequential and structural representations and enhances feature alignment through adversarial learning and contrastive loss, resulting in robust and high-performance DTI prediction.
Jiaqi Deng, Senyu Tang, Ji-Jun Tang et al.· Journal of Chemical Informat...· 0 citations
By holistically integrating atomic, motif, and global fingerprint information via hypergraph modeling, HyperMolFusion offers a more reliable computational tool to enhance the efficiency and accuracy of drug development pipelines.
Yawen Lin, Sheng Lian, Shaoxin Bian et al.· IEEE journal of biomedical a...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.