Jun 2026· International Journal of Molecular Sciences· Vol 27· 0 citations· 44 references
Medicine
TL;DR
HyBind-NN is developed, a multimodal graph neural network that integrates protein language models (PLMs) with 3D structural and dynamic datasets to predict protein–protein and protein–peptide affinity, and it is demonstrated that combining ESM-2 sequence embeddings with precise 3D Voronoi spatial geometry enables accurate affinity predictions across diverse structural datasets.
Abstract
Protein–protein and protein–peptide interactions are fundamental to biological processes, making the accurate prediction of their binding affinity crucial for drug design and mutational analysis. Here, we develop HyBind-NN, a multimodal graph neural network that integrates protein language models (PLMs) with 3D structural and dynamic datasets to predict protein–protein and protein–peptide affinity. First, we demonstrate that combining ESM-2 sequence embeddings with precise 3D Voronoi spatial geometry enables accurate affinity predictions across diverse structural datasets. Next, we show that the inherent limitations of static rigid-body structures can be mitigated through a multi-task learning framework. By utilizing residue-level root mean square fluctuations (RMSF) derived from molecular dynamics (MD) as an auxiliary training target, the model implicitly learns to capture the conformational entropy of flexible peptides without requiring computationally expensive MD simulations during inference. In our benchmarking study, we observe that this multimodal architecture outperforms both purely sequence-based and strictly structural state-of-the-art algorithms, achieving a mean absolute error of 0.89 for pKD (1.12 kcal/mol for ∆G) on the independent benchmark. Finally, we confirmed through ablation analysis that while the PLM provides the dominant predictive signal, geometric representations and dynamic regularization are crucial for resolving subtle conformational rearrangements. This study highlights the synergistic potential of combining PLMs with physics-aware architectures and demonstrates their application towards the robust prediction of intermolecular binding affinity.
GeoPep is introduced, a novel framework for peptide binding site prediction that leverages transfer learning from ESM3, a multimodal protein foundation model that significantly outperforms existing methods in protein–peptide binding site prediction.
Dian Chen, Yunkai Chen, Tong Lin et al.· Journal of Chemical Informat...· 0 citations
Multiple sequence alignment (MSA) Pairformer is presented, a protein language model that builds on AlphaFold2/3's bidirectional refinement between sequence and pairwise residue representations to accurately model the evolution of protein-protein interactions, despite training exclusively on individual chains.
Yo Akiyama, Zhidian Zhang, Olivia Tang et al.· Cell· 2 citations
Mapping the protein interactome is fundamental to understanding disease mechanisms and facilitating therapeutic development. Although protein language models (PLMs) such as ESM-2 have advanced protein-protein interaction (PPI) prediction, their high-dimensional representations remain difficult to connect to verifiable biological signals. To address this limitation, we propose HybridStack-PPI, a gray-box framework that combines ESM-2 sequence representations with explicit physicochemical and motif-derived biological descriptors. The architecture uses motif-anchored local pooling global mean pooling, symmetric pair encoding, fold-internal feature selection, LightGBM branch learners, and an elastic-net logistic-regression stacking layer. We evaluated the method using a C3 cluster-based cross-validation protocol with a 40% sequence-identity clustering threshold and a Same-GO hard-negative setting in which negative candidates shared functional annotations with positive pairs. Under this setting, HybridStack-PPI reached a Human ROC-AUC of 73.65%, PR-AUC of 91.35%, MCC of 28.06%, and specificity of 75.61%. The results indicate a conservative operating point: compared to more recall-oriented baselines, the proposed stack trades lower recall and F1 for higher specificity, MCC, and ranking behavior under functionally similar negative samples. We further reported cross-species transfer, ablation, latency, SHAP-based descriptor attribution, and meta-learner coefficient analyses to clarify both the promise and limitations of biologically informed PPI prediction.
T. T. Nguyen, X. Mai, N. Nguyen· IEEE Access· 0 citations
OrgNet+, a conformational ensemble-aware and orientation-gnostic framework that explicitly incorporates protein structure flexibility during training, is introduced, which substantially reduces intra-ensemble prediction variance while simultaneously improving predictive accuracy.
A. Sarycheva, Aleksandr Shumilov, Petr Popov· Bioinformatics· 0 citations
Abstract Motivation Protein dynamics are central to function, but experiments and molecular dynamics (MD) simulations remain costly, low-throughput, and difficult to compare across protocols. Scalable structure-based methods are needed to infer dynamics from static protein structures. Results We present a deep learning framework that predicts protein dynamics from 30-dimensional Gaussian integral (GI) descriptors of Cα backbone topology. Using 1374 ATLAS protein chains with MD-derived RMSF, GI stratified proteins into fold-relevant clusters enriched for secondary structure, sequence homology, and ECOD families. An attention-based 1D-CNN classified flexible versus non-flexible proteins with test AUC = 0.772 and separated slow-mode– from fast-mode–dominated dynamics with AUC = 0.91. Regression models recovered mean RMSF (Pearson r = 0.72; R² = 0.46) and slow-mode RMSF more accurately (Pearson r = 0.83; R² = 0.62), supporting rapid inference of flexibility and collective-motion bias. Availability and implementation Code and data are available on GitHub at: https://github.com/fvilicich/gaussian_integral/blob/main/gaussian_integral_classification.ipynb.
F. Vilicich, Nicolás Bottino, Zhaoqian Su et al.· Bioinform.· 0 citations
A neural network-based pipeline that integrates amino acid sequences with structural features is developed and provides a modular prototype for follow-up, more extensive protein modeling, including larger proteins and sequence of variable sizes.
Carl David Jasper Causin, M. Fyta· APL Machine Learning· 0 citations