Skip to content
Open access

PMPNN-DDG: an accurate machine learning-based ΔΔG prediction pipeline trained on a novel interpretable feature set extracted from ProteinMPNN

Aug 2026 · bioRxiv · 0 citations · 29 references
Biology

TL;DR

PMPNN-DDG, a Random Forest-based DDG prediction model, trained on a novel set of interpretable features extracted from the recently proposed message-passing neural network-based fixed backbone protein design model, ProteinMPNN, is proposed.

Abstract

An accurate and tractable approximation of the single-point mutation-induced change in protein thermodynamic stability, denoted by DDG, is critical for understanding the genotype– phenotype relationship. Several computational methods have been proposed for this problem; however, limited and error-prone training data and the difficult-to-predict magnitude of structural perturbations make this a challenging task. Consequently, the computational predictors proposed throughout the past decade incrementally improved prediction performance by proposing novel features, combining existing features, task-adapted neural network architectures, loss functions, data augmentation techniques, and pre-training procedures. In this work, we propose PMPNN-DDG, a Random Forest-based DDG prediction model, trained on a novel set of interpretable features extracted from the recently proposed message-passing neural network-based fixed backbone protein design model, ProteinMPNN. On the S669 independent test set, PMPNN-DDG achieves rF +R = 0.64 and RMSE = 1.45, outperforming all compared baseline methods across the reported evaluation measures. On the Ssym independent test set, it achieves rF +R = 0.81, rF−R = −0.99, and RMSE = 1.10, showing competitive performance relative to the compared baselines. PMPNN-DDG is publicly available at https://github.com/dRanger666/PMPNN-DDG.

Read PDF

Similar papers

Aug 2026

A hybrid CNN-GNN-XGB ensemble framework for prediction of mutation-induced protein-protein binding free-energy changes (ΔΔG) in protein engineering.

A hybrid ensemble framework integrating XGBoost, convolutional neural networks (CNN), and graph neural networks (GNN) trained on a curated SKEMPI v2.0 dataset provides a robust and practical tool for ΔΔG prediction with potential applications in protein engineering and rational mutation design.

Sowmya Hari, R. Babu · 0 citations
Open access Jul 2026

T-pGNN4DTI: Towards better drug-target interactions prediction using Global Self-attentive Pooled Graph Convolutional Networks and protein pre-training Models

A DTI prediction method based on the global self-attentive pooled graph neural network and protein pretraining model, called T-pGNN4DTI, which uses a global self-attention pooled graph neural network to learn more meaningful features of the drug molecule.

Yanmei Lin, Bo-Qi Yang, Jianping Liao et al. · 0 citations
Open access Aug 2026

In-Context Learning Meets Small Molecule Property Prediction: Benchmarking Novel Machine Learning Approaches

Recently, a new category of machine learning approaches for tabular data has emerged: tabular foundation models (TFM), based on in-context learning. A TFM is a neural network (usually a transformer) pretrained primarily on synthetic data. Its input is an entire data set: features and labels for training records, along with features for the target records. This study comprehensively benchmarked these methods for predicting the properties of small organic molecules. Several TFM were compared with multiple machine learning methods across 11 data sets (regression, random and structure-aware splits, up to 10,000 molecules each). The evaluation showed that TFM consistently outperform XGBoost, CatBoost, multilayer perceptrons, and other descriptor-based methods, even with careful hyperparameter selection for the baselines. TFM also demonstrate accuracy on par with or better than graph-based methods, including those pretrained on chemical data. Uni-Mol2, a pretrained deep neural network operating on 3D atomic coordinates, slightly outperforms TFM in some experiments. However, this comparison deliberately disfavors TFM, as they do not use any chemical pretraining and rely on a minimalistic set of 2D molecular descriptors without feature engineering. Some further improvement in results for relatively large data sets and random splits is achieved using retrieval: for each test molecule, the 500 closest neighbors (by Tanimoto similarity) are selected from the training set, and TFM inference is performed on this local subset. Overall, TFM (particularly the TabPFN family) are highly promising for predicting the properties of small molecules.

Unknown authors · 0 citations
Aug 2026

To ML-Predict or Not to ML-Predict: The Impact of Machine Learning-Predicted Protein Structures on FEP Accuracy and Data Augmentation

The rapid advancement of machine learning (ML)-based protein structure prediction, exemplified by AlphaFold2 and extended by newer models such as AlphaFold3 and Boltz-2, has generated significant optimism for structure-guided drug discovery. In particular, ligand–protein cofolding approaches offer the potential to overcome limitations in generating starting structures for physics-based free energy perturbation (FEP) calculations. However, the practical readiness of ML-predicted structures for FEP applications remains insufficiently evaluated. Here, we systematically assess experimentally determined crystal structures, a homology model, and ML-predicted protein structures as inputs for FEP using a well-characterized congeneric series targeting the tyrosine kinase cSrc. A data set of 133 compounds was evaluated through more than 1400 FEP calculations under minimal optimization to approximate “out-of-the-box” performance. By maintaining consistent preparation protocols, we isolate the impact of structural origin on predictive accuracy. Variable performance was observed across both experimental and ML-predicted structures, highlighting that even under this idealized benchmark scenario, significant challenges remain in reliably generating and refining predictive protein–ligand complexes. This study demonstrates that predictive variation in micro and macro conformational states─rather than the structural source─governs predictive reliability, underscoring the need for careful validation when integrating ML-derived structures into FEP workflows.

Parker Dryja, Morné Muller, Monique Horn et al. · 0 citations
Aug 2026

SA-MPNN: A Sequence-Aware ThermoMPNN for Accurate Prediction of Mutational Effects on Protein Thermodynamic Stability

Predicting the impact of single-point mutations on protein thermodynamic stability is crucial for protein engineering of therapeutic and industrial applications. By effectively capturing the three-dimensional structural information of proteins and the spatial physical environment of each residue, the inverse folding models (IFMs) upon fine-tuning, such as ThermoMPNN, achieved state-of-the-art performance in predicting thermostability changes in proteins caused by mutations. However, IFMs are limited in their capacity to capture protein deep evolutionary information, whereas protein language models (pLMs) excel. Here, we present SA-MPNN, a lightweight, end-to-end hybrid framework that dynamically integrates the protein sequence representations from a protein language model (ESM2) into the ThermoMPNN architecture to improve protein stability prediction by combining evolutionary representations with geometric structural embeddings. By evaluating various feature fusion strategies, we selected a self-attention-based integration mechanism to effectively combine the two modalities. Trained on the large-scale Megascale data set, SA-MPNN achieved modest but consistent gains over ThermoMPNN on various benchmark data sets, with particularly noticeable improvements in several correlation analysis and screening-oriented evaluations. Finally, wet-lab validation was performed on the top-ranking variants of Acetivibrio thermocellusβ-glucosidase (AtBgl1A) as a case study. The experimental results demonstrated that multiple designed mutants exhibited improved thermostability, and the optimal variant, GC20, achieved a melting temperature (Tm) of 76.98 °C, representing a 5.97 °C increase over the wild-type, thereby supporting the practical applicability of SA-MPNN in protein engineering.

Xin-Yue Zhang, Xiang Zheng, Ze-Yuan Dong et al. · 0 citations
Jul 2026

Meta-learning GNN with MD-informed attention for cross-species prediction of phosphoinositide-dependent kinase-1 (PdK1) inhibitors in termite control

This work bridges the gap between data-driven and physics-based approaches, providing a scalable solution for pesticide discovery when target-specific data are limited, and incorporating meta-learning and MD insights improves cross-species transferability and yields interpretable attention patterns based on biophysical principles.

Haroon · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.