Jun 2026· Journal of Chemical Information and Modeling· Vol 66, pp. 7377 - 7389· 0 citations· 50 references
Medicine
TL;DR
GeoPep is introduced, a novel framework for peptide binding site prediction that leverages transfer learning from ESM3, a multimodal protein foundation model that significantly outperforms existing methods in protein–peptide binding site prediction.
Abstract
Multimodal approaches that integrate protein structure and sequence have achieved remarkable success in protein–protein interface prediction. However, extending these methods to protein-peptide interactions remains challenging due to the inherent conformational flexibility of peptides and the limited availability of structural data that hinders direct training of structure-aware models. To address these limitations, we introduce GeoPep, a novel framework for peptide binding site prediction that leverages transfer learning from ESM3, a multimodal protein foundation model. GeoPep fine-tunes ESM3′s rich prelearned representations from protein–protein binding to address the limited availability of protein–peptide binding data. The fine-tuned model is further integrated with a Kolmogorov–Arnold Network (KAN)-based architecture for complex nonlinear approximation. Furthermore, the model is trained using distance-based loss functions that exploit 3D structural information to enhance binding site prediction. Comprehensive evaluations demonstrate that GeoPep significantly outperforms existing methods in protein–peptide binding site prediction by effectively capturing sparse and heterogeneous binding patterns.
HyBind-NN is developed, a multimodal graph neural network that integrates protein language models (PLMs) with 3D structural and dynamic datasets to predict protein–protein and protein–peptide affinity, and it is demonstrated that combining ESM-2 sequence embeddings with precise 3D Voronoi spatial geometry enables accurate affinity predictions across diverse structural datasets.
E. A. Bogdanova, A. Chernukhin, Alexey K. Shaytan· International Journal of Mol...· 0 citations
Computational modeling provides geometric insight into protein-protein interactions without requiring the resources of experimentation. However, reliability can be hindered when modeling proteins with distinctive features, such as antibodies, that use flexible, polar-rich loops to bind antigens. We developed ProteinDock, a physics-based tool that can be used in combination with leading modeling programs to improve the reliability of protein-protein docking; this work provides a case study of antibody-antigen interfaces. ProteinDock was layered onto Rosetta for docking unbound experimentally determined structures, and when evaluated on Docking Benchmark Set 5.5, generated CAPRI acceptable-quality or better for 80.2% of targets, an improvement of 32.8 percentage points over vanilla Rosetta’s 47.4% on the same dataset. To improve protein-protein prediction reliability from sequence inputs, we demonstrate that a truncated version of ProteinDock can be used to choose the optimal prediction among outputs from multiple deep learning-based tools. We show that this strategy is a computationally efficient alternative to increasing the seed quantity for deep-learning predictions. A graphical user interface for layering ProteinDock has been created and is available at https://github.com/Kimmel-Lab/proteindock and https://proteindock.com/. TOC Figure
G. Rajagopal, Søren C. Spina, Joe Bailey et al.· bioRxiv· 0 citations
We present PepEDiff, a novel peptide binder generator that designs binding sequences given a receptor protein and the target pocket residues. Peptide binder generation is critical in therapeutic and biochemical applications, yet many existing methods rely heavily on intermediate structure prediction, adding complexity and limiting sequence diversity. Our approach departs from this paradigm by generating binder sequences directly in a continuous latent space derived from a pretrained protein embedding model, without relying on predicted structures, thereby improving structural and sequence diversity. To encourage the model to capture binding-relevant features rather than memorizing known sequences, we perform latent-space exploration and diffusion-based sampling, enabling the generation of peptides beyond the limited distribution of known binders. This out-of-distribution generative strategy leverages the global protein embedding manifold as a semantic prior, allowing the model to propose novel peptide sequences in previously unseen regions of the protein space. We evaluate PepEDiff on TIGIT, a challenging target with a large, flat protein–protein interaction interface that lacks a druggable pocket. Despite its simplicity, our method outperforms state-of-the-art approaches across benchmark tests and in the TIGIT case study, demonstrating its potential as a general, structure-free framework for zero-shot peptide binder design. The code for this research is available at https://anonymous.4open.science/r/PepEDiff-/
Po-Yu Liang, Tibo Duran, Jun Bai· ACM International Conference...· 1 citation
Protein phosphorylation regulates signaling, yet atomic-level substrate specificity remains elusive due to sparse structural data and phosphorylation-site-insensitive deep-learning predictors. Here we present a pipeline reformulating kinase-substrate modeling as a Bayesian inference problem. By integrating curated data sets and literature evidence parsed by Large Language Models, we converted diverse biological knowledge into structural restraints for the restraint-guided deep-learning model GRASP. For EGFR, BRAF and JNK1, we obtained 336 new phosphorylation-site-specific structure candidates refined by molecular dynamics. These models recapitulate known features, such as JNK1's hydrophobic docking groove, and enabled a Virtual Position Scanning Peptide Array (V-PSPA) to map recognition patches and derive sequence preferences. Cross-referencing predicted interfaces with AlphaMissense pathogenicity scores reveal that the interaction types and distances to the catalytic pocket significantly influence pathogenicity scores. A comparison with clinical mutation data sets further connects pathogenic mutations to the kinase-substrate interface. This high-resolution, high-throughput pipeline can be broadly applicable to kinase specificity studies and general drug discovery.
Jinyuan Hu, Shimian Li, Yue Xue et al.· Journal of Chemical Informat...· 0 citations
The functional classification of protein sequences remains a major bottleneck in biology. Although protein language model (PLM)-based approaches have substantially improved broad protein function prediction, most protein sequences still lack precise annotation at the level of specialized functions—the fine-grained molecular roles that define specificity within protein families. We present PLMView, an unsupervised framework for fine-grained protein function classification directly from sequence. PLMView reframes protein function inference as a relational problem: instead of embedding sequences in isolation, it positions them within a collaborative functional space defined by comparisons with PLM embeddings of anchor sequences, thereby capturing subtle sequence–function relationships. Without requiring labeled data, family-specific training, or PLM fine-tuning, PLMView accurately distinguishes specialized functions among homologous proteins and highlights residues likely to determine functional specificity. The method achieves high precision while remaining computationally efficient, classifying approximately 10,000 sequences with 1,000 anchors in under 40 minutes; compared with pooled-embedding approaches and, in challenging cases, Sequence Similarity Networks, PLMView provides finer and more biologically coherent functional resolution, while achieving more than 10-fold speed-up over SSN reconstruction on datasets of this scale. Applications to thioredoxins, visual opsins, and Tara Oceans environmental diatom cold-shock proteins show that PLMView can move from interpretable residue-level determinants in well-studied protein families to large-scale environmental functional discovery, linking molecular specialization to ecological distribution and transcriptional deployment across the global ocean.