Aug 2025· Journal of Chemical Information and Modeling· Vol 65, pp. 9061-9074· 4 citations· 37 references
MedicineComputer Science
TL;DR
PepBAN is introduced, a deep learning framework for modeling PepPI predictions that effectively learns the pattern of pairwise local interactions, enables the identification of key residues participating in the peptide-protein interactions, and offers an intuitive approach to interpret the underlying mechanisms of PepPIs via analyzing attention weights.
Abstract
Accurate prediction of the peptide-protein interaction (PepPI) is crucial for developing peptide-based therapeutics and vaccines. However, this computational task has traditionally faced significant challenges, such as the scarcity of structure data along with the corresponding label of the binding affinity for bound complexes. To address these challenges, we introduce PepBAN, a deep learning framework for modeling PepPI predictions. PepBAN incorporates two technical advancements: (1) adopting the protein language model ESM-2 to characterize proteins and ESM-2 or a graph-based foundation model for peptides without structure data and (2) leveraging the conditional domain adversarial learning to enhance generalization across a broad range of protein targets, especially when there are limited binding data. At the core of PepBAN is a bilinear attention network (BAN) that effectively learns the pattern of pairwise local interactions, enables the identification of key residues participating in the peptide-protein interactions, and offers an intuitive approach to interpret the underlying mechanisms of PepPIs via analyzing attention weights. Our numerical experiments demonstrated that PepBAN outperformed the previous state-of-the-art models across several well-established benchmark studies. Furthermore, we evaluated PepBAN's applicability in predicting cyclic peptide-protein interactions, a task that poses significant challenges due to the presence of noncanonical amino acids. These nonstandard residues require specialized handling, which most existing sequence-based PepPI prediction models did not adequately address, and we adopt an atom-resolved molecular graph approach to process cyclic peptides. Despite this complexity, PepBAN demonstrated a clear advantage by achieving a superior prediction performance and offering a distinct edge in tackling the emerging chemical space of cyclic peptides, which has great potential for novel therapeutic development. In summary, PepBAN serves as a valuable tool for advancing peptide-based drug and therapeutic development.
Experimental results on multiple benchmark data sets demonstrate that MMU-DPI outperforms several state-of-the-art DPI prediction methods and indicate that MMU-DPI can serve as a useful computational tool for drug discovery.
Jiahao Wei, Tie Shen· Journal of Chemical Informat...· 0 citations
The ImmunoFoundation Model (IFM), a multimodal deep learning system that integrates not only peptide sequences, 3D molecular structures, and biochemical properties but also TCR-MHC-peptide to achieve superior immunogenicity prediction and enable peptide optimization for therapeutic applications is developed.
Smita Krishnaswamy, J. Rocha, Hiren Madhu et al.· Journal of Immunology· 0 citations
CLDN18.2 is a promising tumor-specific antigen; however, the development of therapeutic antibodies against it is challenged by the need for simultaneous optimization of affinity and developability. To address this, we present cdrGPT, a deep learning framework based on GPT-2 for de novo generation of complementarity-determining region H3 (CDRH3) sequences. Our approach integrates pre-training on the Observed Antibody Space (OAS) database with structural templating derived from the known antibody zolbetuximab. Generated sequences were iteratively refined through rejection sampling and fine-tuned against a multi-parameter objective function encompassing predicted affinity and MHC class II binding risk. From an initial set of 50,000 sequences, this screening pipeline yielded 313 high-confidence candidates. Subsequent analysis using evolutionary scale modeling 2 (ESM2) embeddings, principal component analysis (PCA), and clustering revealed three structurally distinct clusters, with intra-cluster cosine similarities exceeding 0.99. Validation of seven representative sequences from the dominant cluster using AlphaFold3 confirmed high structural fidelity to the zolbetuximab template, demonstrating a root mean square deviation (RMSD) of 1.331 Å for the CDRH3 loop and positional deviations of less than 0.4 Å for key paratope residues. These results indicate that the designed variants preserve the core binding mode of the parent antibody. This study establishes a feasible pipeline for integrating AI-generated CDRH3 loops into functional antibody scaffolds, providing a foundation for the accelerated development of therapeutics targeting CLDN18.2 and other clinically relevant antigens.
Tao Qu, Lingyan Yuan, Wei-Ran Cui et al.· PLoS Computational Biology· 0 citations
This work designs a biological prior-guided feature fusion framework that integrates pseudo-structural epitope knowledge and CDR-specific attention mechanisms via a mixture-of-experts architecture to effectively capture complex binding landscapes in antibody screening and drug residence time analysis.
G. Luo, Junkai Wang, Sizhe Zhang et al.· Bioinformatics· 0 citations
Accurate computational prediction of antiviral peptides (AVPs) can accelerate peptide screening and reduce experimental costs. However, existing deep learning-based methods still suffer from severe class imbalance, over-reliance on handcrafted features and limited interpretability. Here, we propose PMAVP, a multi-task learning framework that integrates the ProtT5 pre-trained protein language model with a Mamba-inspired module for AVP identification and functional activity prediction. We use ProtT5 to extract deep semantic representations from peptide sequences and a Mamba module to capture long-range dependencies at a lower computational complexity. We introduce Focal Loss to mitigate class imbalance and leverage transfer learning to enhance performance on functional activity prediction. Experimental results demonstrate that our model achieves superior performance in terms of prediction accuracy, stability, and computational efficiency. Furthermore, DeepSHAP-based interpretability analysis reveals that the first 40 amino acid residues contribute substantially to AVP prediction.
Pei-Wei Wei, Wei-Hao Su, Qing-Song Qin et al.· International Journal of Mol...· 0 citations
This paper proposes TextDTI, a multimodal framework that simultaneously exploits sequential and structural representations and enhances feature alignment through adversarial learning and contrastive loss, resulting in robust and high-performance DTI prediction.
Jiaqi Deng, Senyu Tang, Ji-Jun Tang et al.· Journal of Chemical Informat...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 18, 2026
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.
What does it take to trust AI-driven HVAC optimization? Our AI Model Factory combines agents, machine learning, reinforcement learning and deterministic checks in a governed workflow designed for messy, real-world building data. The post We built an AI factory for HVAC control appeared first on GPT-Lab.
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
MIT News · Artificial Intelligence· news.mit.eduAug 3, 2026