Aug 2026· ACS Physical Chemistry Au· 0 citations· 39 references
Abstract
Predicting protein–protein binding free energy (ΔG) from structure remains a central challenge in computational biophysics. Here, we present GULP (Graph-based Unified Learning for Protein binding), a graph neural network (GNN) that jointly learns from a residue-level graph representation of the binding interface and global physicochemical descriptors. We systematically investigate how training data distribution affects model performance by comparing a full training set with a balanced subset enriched for extreme-affinity complexes. GULP is computationally efficient and provides interpretable insights into residue-level and physicochemical contributions to binding. On external validation, GULP achieves a mean absolute error (MAE) of 2.31 kcal/mol and shows moderate agreement with experimental ΔG values (Pearson r = 0.54, Spearman ρ = 0.58).
HyBind-NN is developed, a multimodal graph neural network that integrates protein language models (PLMs) with 3D structural and dynamic datasets to predict protein–protein and protein–peptide affinity, and it is demonstrated that combining ESM-2 sequence embeddings with precise 3D Voronoi spatial geometry enables accurate affinity predictions across diverse structural datasets.
E. A. Bogdanova, A. Chernukhin, Alexey K. Shaytan· International Journal of Mol...· 0 citations
Accurate identification of near-native ligand binding poses is a central challenge in structure-based drug design. From a physical point of view, the successful construction of a protein–ligand complex structure is dependent on whether protein and ligand can form enough atomically pairwise interactions that result in a global energy minimum. In this work, we report a machine learning scoring strategy for protein–ligand screening which explicitly considers the Native Contact Ratio (NCR), a topology inspired metric that quantifies the preservation of protein–ligand interfacial contacts as well as interaction energy. This physics-awared supervision strategy provides a simple but efficient gradient field that faithfully reflects the complicated protein energy landscape than conventional 3D coordinate-based objectives. Building on this principle, we present DeepNCR, an energy-informed Transformer framework that encodes approximate Coulombic and dispersive interaction potentials across the protein–ligand binding interface. Furthermore, we introduce a feature pruning step that compresses the interaction tensor from 1470 to 868 dimensions, further improving signal-to-noise ratio and directing model attention toward the interaction motifs critical for binding specificity. The model optimizes topological objectives and at inference drives pose refinement through a differentiable hybrid gradient field integrating predicted NCR and AutoDock Vina energetics. Extensive evaluation on the CASF-2016 benchmark and the 3D-DISCO cross-docking data set demonstrates consistently high performance: a Top-1 docking success rate of 94.7%, a 1% Enrichment Factor of 21.21 in virtual screening, and a Top-1 cross-docking success rate of 34.8%. Mechanistic analysis reveals that NCR-guided optimization enables decoy escaping from local energy minima and drives the recovery of disrupted native interactions, confirming that NCR captures the physical determinants of binding rather than mere geometric proximity.
Zhenqiang Zhang, Zhihao Wang, Yang Liu et al.· Journal of Chemical Informat...· 0 citations
Abstract Motivation Protein dynamics are central to function, but experiments and molecular dynamics (MD) simulations remain costly, low-throughput, and difficult to compare across protocols. Scalable structure-based methods are needed to infer dynamics from static protein structures. Results We present a deep learning framework that predicts protein dynamics from 30-dimensional Gaussian integral (GI) descriptors of Cα backbone topology. Using 1374 ATLAS protein chains with MD-derived RMSF, GI stratified proteins into fold-relevant clusters enriched for secondary structure, sequence homology, and ECOD families. An attention-based 1D-CNN classified flexible versus non-flexible proteins with test AUC = 0.772 and separated slow-mode– from fast-mode–dominated dynamics with AUC = 0.91. Regression models recovered mean RMSF (Pearson r = 0.72; R² = 0.46) and slow-mode RMSF more accurately (Pearson r = 0.83; R² = 0.62), supporting rapid inference of flexibility and collective-motion bias. Availability and implementation Code and data are available on GitHub at: https://github.com/fvilicich/gaussian_integral/blob/main/gaussian_integral_classification.ipynb.
F. Vilicich, Nicolás Bottino, Zhaoqian Su et al.· Bioinform.· 0 citations
Quantifying protein–protein binding affinity is essential for understanding molecular recognition and guiding antibody and inhibitor design. However, binding affinity is governed by tightly coupled sequence, structural, and chemical determinants. Existing models often encode these factors in isolation, limiting their ability to capture the multi-level dependencies underlying binding affinity
$$\left(\Delta\text{G}\right)$$
.
We propose MIRAGE, a graph-based framework for direct
$$\Delta \text{G}$$
prediction that explicitly models interactions across multidimensional (1D sequences, 2D contact maps, 3D structures) and multi-scale (residue-level, atom-level) features. MIRAGE integrates two complementary modules to capture cross-dimensional and cross-scale dependencies, enabling unified residue–atom representation learning. Across public benchmarks, MIRAGE demonstrated strong generalization, achieving Pearson correlations of 0.70 and 0.69 on two independent external test sets and retaining predictive effectiveness under structure-separated cross-validation designed to reduce structural information sharing. In a supplementary analysis with AlphaFold3-predicted complex structures, MIRAGE also preserved significant predictive correlations when experimentally resolved structures were unavailable. Ablation studies confirm the contributions of each module. Interpretability analyses further show that the model focuses on biophysically meaningful interface regions. The source code of MIRAGE is available from
https://github.com/ShiweiWu-545/MIRAGE
.
These results indicate that explicitly modeling multi-level interactions is important for accurately capturing the determinants of binding affinity. MIRAGE provides an interpretable and robust framework for structure-aware
$$\Delta \text{G}$$
prediction, with potential applications in protein engineering and drug design.
Predicting favorable protein-peptide binding events remains a central challenge in biophysics, with continued uncertainty surrounding how nonlocal effects shape the global energy landscape. Here, we introduce peripheral surface information (PSI) entropy, SΨ, a quantitative measure of the statistical variability in apolar and charged non-interacting surface (NIS) proportions across conformational ensembles. Within the Gibbs free-energy relation ΔG = ΔH - TΔS, SΨ is proposed as a computationally tractable entropic proxy rather than a direct thermodynamic observable or stand-alone estimator of binding affinity. Using energy-directed molecular docking via HADDOCK3 and explicit-solvent molecular dynamics simulations, it is demonstrated that favorable binding partners exhibit emergent, low-entropy N-states (discrete macrostates in NIS state space) indicative of preferential apolar/charged surface configurations. Across dozens of peptides and multiple receptor systems (WW, PDZ, and MDM2 domains), dominant N-states persisted under varied docking parameters and initial conditions. A meta-ensemble of 657 complexes from 36 experiments over 15 years confirmed the presence of dominant NIS modes independent of in silico methodology, suggesting an evolutionary selection pressure toward specific NIS fingerprints. These findings establish SΨ as a thermoinformatic descriptor that encodes favorable binding constraints into unique statistical signatures of the NIS.
Tyler Grear, Donald J. Jacobs· Biophysical Journal· 0 citations
Accurate prediction of protein-protein interaction interfaces is critical for understanding molecular recognition and guiding therapeutic design. This study presents a comprehensive machine learning pipeline for predicting interface residues in permanent homodimeric protein complexes. Using a curated dataset of 1311 homodimers, we benchmarked six widely used machine learning algorithms and identified multilayer perceptron and XGBoost as top performers, achieving Matthews correlation coefficients (MCC) exceeding 0.93. To enhance interpretability and efficiency, we employed recursive feature elimination to derive a minimal set of six biologically meaningful features, including solvent accessibility, surface roughness, planarity, and average protrusion index, that retained high predictive power (MCC > 0.90). Structurally stratified models tailored to α-helical, β-strand, and membrane proteins demonstrated comparable or improved accuracy relative to generalized models, particularly when utilizing the reduced feature subset. As a preliminary demonstration of generalizability, we applied our approach to an external heterodimer complex (PDB ID: 9ETL). While limited to a single case study, the structurally specialized models maintained high accuracy, suggesting potential applicability beyond the training domain. Furthermore, our residue-level feature-driven models demonstrated highly competitive performance when compared against the baseline established by the general-purpose ColabFold pipeline. The results highlight the importance of structural context in interface prediction and demonstrate that compact, structure-aware models can achieve high accuracy while reducing computational complexity. This work provides a scalable, interpretable, and biologically informed approach to protein interface prediction, with implications for large-scale structural descriptor, drug target characterization, and protein engineering applications.
Tayyip Topuz, Z. Erdem, Halil Bisgin et al.· Scientific Reports· 0 citations