PMPNN-DDG, a Random Forest-based DDG prediction model, trained on a novel set of interpretable features extracted from the recently proposed message-passing neural network-based fixed backbone protein design model, ProteinMPNN, is proposed.
Abstract
An accurate and tractable approximation of the single-point mutation-induced change in protein thermodynamic stability, denoted by DDG, is critical for understanding the genotype– phenotype relationship. Several computational methods have been proposed for this problem; however, limited and error-prone training data and the difficult-to-predict magnitude of structural perturbations make this a challenging task. Consequently, the computational predictors proposed throughout the past decade incrementally improved prediction performance by proposing novel features, combining existing features, task-adapted neural network architectures, loss functions, data augmentation techniques, and pre-training procedures. In this work, we propose PMPNN-DDG, a Random Forest-based DDG prediction model, trained on a novel set of interpretable features extracted from the recently proposed message-passing neural network-based fixed backbone protein design model, ProteinMPNN. On the S669 independent test set, PMPNN-DDG achieves rF +R = 0.64 and RMSE = 1.45, outperforming all compared baseline methods across the reported evaluation measures. On the Ssym independent test set, it achieves rF +R = 0.81, rF−R = −0.99, and RMSE = 1.10, showing competitive performance relative to the compared baselines. PMPNN-DDG is publicly available at https://github.com/dRanger666/PMPNN-DDG.
A hybrid ensemble framework integrating XGBoost, convolutional neural networks (CNN), and graph neural networks (GNN) trained on a curated SKEMPI v2.0 dataset provides a robust and practical tool for ΔΔG prediction with potential applications in protein engineering and rational mutation design.
Sowmya Hari, R. Babu· Computational biology and ch...· 0 citations
A DTI prediction method based on the global self-attentive pooled graph neural network and protein pretraining model, called T-pGNN4DTI, which uses a global self-attention pooled graph neural network to learn more meaningful features of the drug molecule.
Recently, a new category of machine learning approaches for tabular data has emerged: tabular foundation models (TFM), based on in-context learning. A TFM is a neural network (usually a transformer) pretrained primarily on synthetic data. Its input is an entire data set: features and labels for training records, along with features for the target records. This study comprehensively benchmarked these methods for predicting the properties of small organic molecules. Several TFM were compared with multiple machine learning methods across 11 data sets (regression, random and structure-aware splits, up to 10,000 molecules each). The evaluation showed that TFM consistently outperform XGBoost, CatBoost, multilayer perceptrons, and other descriptor-based methods, even with careful hyperparameter selection for the baselines. TFM also demonstrate accuracy on par with or better than graph-based methods, including those pretrained on chemical data. Uni-Mol2, a pretrained deep neural network operating on 3D atomic coordinates, slightly outperforms TFM in some experiments. However, this comparison deliberately disfavors TFM, as they do not use any chemical pretraining and rely on a minimalistic set of 2D molecular descriptors without feature engineering. Some further improvement in results for relatively large data sets and random splits is achieved using retrieval: for each test molecule, the 500 closest neighbors (by Tanimoto similarity) are selected from the training set, and TFM inference is performed on this local subset. Overall, TFM (particularly the TabPFN family) are highly promising for predicting the properties of small molecules.
Unknown authors· Journal of Chemical Informat...· 0 citations
The rapid advancement of machine learning (ML)-based protein structure prediction, exemplified by AlphaFold2 and extended by newer models such as AlphaFold3 and Boltz-2, has generated significant optimism for structure-guided drug discovery. In particular, ligand–protein cofolding approaches offer the potential to overcome limitations in generating starting structures for physics-based free energy perturbation (FEP) calculations. However, the practical readiness of ML-predicted structures for FEP applications remains insufficiently evaluated. Here, we systematically assess experimentally determined crystal structures, a homology model, and ML-predicted protein structures as inputs for FEP using a well-characterized congeneric series targeting the tyrosine kinase cSrc. A data set of 133 compounds was evaluated through more than 1400 FEP calculations under minimal optimization to approximate “out-of-the-box” performance. By maintaining consistent preparation protocols, we isolate the impact of structural origin on predictive accuracy. Variable performance was observed across both experimental and ML-predicted structures, highlighting that even under this idealized benchmark scenario, significant challenges remain in reliably generating and refining predictive protein–ligand complexes. This study demonstrates that predictive variation in micro and macro conformational states─rather than the structural source─governs predictive reliability, underscoring the need for careful validation when integrating ML-derived structures into FEP workflows.
Parker Dryja, Morné Muller, Monique Horn et al.· Journal of Chemical Informat...· 0 citations
Predicting the impact of single-point mutations on protein thermodynamic stability is crucial for protein engineering of therapeutic and industrial applications. By effectively capturing the three-dimensional structural information of proteins and the spatial physical environment of each residue, the inverse folding models (IFMs) upon fine-tuning, such as ThermoMPNN, achieved state-of-the-art performance in predicting thermostability changes in proteins caused by mutations. However, IFMs are limited in their capacity to capture protein deep evolutionary information, whereas protein language models (pLMs) excel. Here, we present SA-MPNN, a lightweight, end-to-end hybrid framework that dynamically integrates the protein sequence representations from a protein language model (ESM2) into the ThermoMPNN architecture to improve protein stability prediction by combining evolutionary representations with geometric structural embeddings. By evaluating various feature fusion strategies, we selected a self-attention-based integration mechanism to effectively combine the two modalities. Trained on the large-scale Megascale data set, SA-MPNN achieved modest but consistent gains over ThermoMPNN on various benchmark data sets, with particularly noticeable improvements in several correlation analysis and screening-oriented evaluations. Finally, wet-lab validation was performed on the top-ranking variants of Acetivibrio thermocellusβ-glucosidase (AtBgl1A) as a case study. The experimental results demonstrated that multiple designed mutants exhibited improved thermostability, and the optimal variant, GC20, achieved a melting temperature (Tm) of 76.98 °C, representing a 5.97 °C increase over the wild-type, thereby supporting the practical applicability of SA-MPNN in protein engineering.
Xin-Yue Zhang, Xiang Zheng, Ze-Yuan Dong et al.· Journal of Chemical Informat...· 0 citations
This work bridges the gap between data-driven and physics-based approaches, providing a scalable solution for pesticide discovery when target-specific data are limited, and incorporating meta-learning and MD insights improves cross-species transferability and yields interpretable attention patterns based on biophysical principles.
Haroon· Journal of Molecular Modelin...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.