Jul 2026· Protein Science· Vol 35· 0 citations· 43 references
Medicine
TL;DR
The proposed RINAMI model is established as an accurate and interpretable framework for ΔG prediction and provide a practical computational tool for evaluating and prioritizing protein designs prior to experimental testing.
Abstract
Recent advances in de novo protein design have enabled the generation of diverse novel proteins. However, a fundamental challenge remains: even when an amino acid sequence is designed with the target structure as the most stable conformation, there is currently no reliable computational method for assessing whether the target structure is sufficiently stabilized relative to alternative conformations. While experimental realization of the intended fold requires the target structure to be thermodynamically favored by a large free‐energy gap, the absence of a quantitative measure of folding stability makes it difficult to distinguish reliable from unreliable designs. Here, we propose the Residue‐attributed Interpretable Neural network for predicting Absolute folding free energy by Merging structure and sequence Information (RINAMI), a machine learning model that predicts the absolute folding free energy (ΔG) of proteins from their three‐dimensional structures and amino acid sequences. RINAMI integrates structure‐ and sequence‐based representations derived from ProteinMPNN and Evolutionary Scale Modeling 2 (ESM2) using a multi‐head cross‐attention mechanism that contextualizes sequence‐derived signals within the structural environment. Benchmarking RINAMI on both natural and designed proteins from the Mega‐scale and Maxwell datasets shows that it outperforms the tested existing approaches, achieving higher correlations with experimental measurements and improved or comparable prediction errors. An ablation study supports the contribution of sequence–structure integration for predictive accuracy. In addition, RINAMI exhibits strong interpretability by capturing key physicochemical effects, including the destabilizing effect of buried hydrophilic residues, the stabilizing effect of buried hydrophobic residues, and the characteristics of cysteine. Together, these results establish RINAMI as an accurate and interpretable framework for ΔG prediction and provide a practical computational tool for evaluating and prioritizing protein designs prior to experimental testing.
This framework provides a clearer understanding of how methodological shifts have shaped the capabilities, limitations, and practical roles of recent models.
Wengan He, Yongsheng Luo, Lihong Jiang et al.· 0 citations
Protein inverse folding aims to recover amino acid sequences for a given 3D protein structure, underpinning broad applications such as enzyme engineering and drug discovery.Current methods often follow a serial pipeline, in which a structure encoder predicts a coarse sequence, which is then refined by protein language models (PLMs). However, because PLMs only perform post-hoc sequence edits, the refinement is bounded by the quality of upstream predictions.Thanks to recent multimodal protein language models (MPLMs), we could directly encode structure to generate sequences with pretrained structural knowledge, but we observe that they are not effective for inverse folding. Therefore, we introduce a symmetric dual-path architecture that both leverages PLMs for pretrained sequence evolution knowledge and MPLMs for pretrained structural knowledge to iteratively guide protein sequence generation.Through extensive experiments across standard protein inverse folding benchmarks, our method achieves state-of-the-art performance, surpassing prior approaches, and ablation studies validate the rationale of our symmetric design, revealing a promising direction for the community.
Han-Dong Wang, Jiaxin Qi, Baisheng Lai et al.· 0 citations
The (un)folding rates of natural proteins determine their native stability and functional homeostasis, making them important targets for protein engineering and design. From a prediction standpoint, the rates have been a long‐standing puzzle. We have known for decades that folding rates empirically correlate with properties of the native three dimensional (3D) structures and that both, folding and unfolding rates, scale with protein size. Whereas such rate correlations are too rough for being of practical use, no significant progress in prediction accuracy has occurred since then, despite many efforts even including machine learning approaches. Here, we retake on this challenge by expanding the simple one‐dimensional free energy surface (1D‐FES) model that originally led to demonstrate the size scaling of both rates, and a curated database with rates for 75 single‐domain proteins. We define the weighted sequence order (WSO) as a novel parameter that allows incorporating structural information into the 1D‐FES model explicitly. Via the WSO, we examine the role of global structural properties such as fold topology and core packing in defining the (un)folding rates within the context of a physics‐based model of protein folding. After introducing fold topology and packing at a coarse‐grained level, the model uses three floating parameters to predict the folding and unfolding rates within 6.5‐ and 10‐fold, respectively, resulting in ±6.5 kJ/mol accuracy in native stability, equivalent to the typical perturbation induced by one single‐point mutation. The net improvement over the 2‐parameter size‐only prediction is of 2.5‐fold. These new rate predictions are significantly closer to the threshold of usefulness for engineering and design. More importantly, this WSO‐modified 1D‐FES model can now directly accommodate atomistic, high‐resolution, force‐fields to further optimize the rate predictions, and/or to use rate information as a testbed for force‐field refinement. Finally, the WSO‐1D‐FES model could also serve as foundation for developing more complex models capable of dealing with multi‐domain proteins as well as with the evolutionary information cryptically encoded in natural protein sequences.
Mohammad Abdulqader, Victor Muñoz· Protein Science· 0 citations
Recent advances in protein structure prediction, exemplified by AlphaFold, have largely addressed the determination of static structures, one aspect of the protein folding problem. However, predicting folding pathways, by which proteins reach their native states, remains a significant challenge. Here, we present PathFold, a deep learning framework that predicts protein folding pathways directly from sequence information. PathFold leverages an AlphaFold-based module to extract structural information from the sequence and generates a progressive folding trajectory from an extended conformation using a diffusion model. By modeling the full trajectory, it enables prediction of folding intermediates and transition pathways, analogous to those observed in steered molecular dynamics (SMD) simulations. The predicted pathways reveal well-defined intermediates and sequential folding events, and show agreement with experimental folding data, including measured Φ-values.
Predicting the impact of single-point mutations on protein thermodynamic stability is crucial for protein engineering of therapeutic and industrial applications. By effectively capturing the three-dimensional structural information of proteins and the spatial physical environment of each residue, the inverse folding models (IFMs) upon fine-tuning, such as ThermoMPNN, achieved state-of-the-art performance in predicting thermostability changes in proteins caused by mutations. However, IFMs are limited in their capacity to capture protein deep evolutionary information, whereas protein language models (pLMs) excel. Here, we present SA-MPNN, a lightweight, end-to-end hybrid framework that dynamically integrates the protein sequence representations from a protein language model (ESM2) into the ThermoMPNN architecture to improve protein stability prediction by combining evolutionary representations with geometric structural embeddings. By evaluating various feature fusion strategies, we selected a self-attention-based integration mechanism to effectively combine the two modalities. Trained on the large-scale Megascale data set, SA-MPNN achieved modest but consistent gains over ThermoMPNN on various benchmark data sets, with particularly noticeable improvements in several correlation analysis and screening-oriented evaluations. Finally, wet-lab validation was performed on the top-ranking variants of Acetivibrio thermocellusβ-glucosidase (AtBgl1A) as a case study. The experimental results demonstrated that multiple designed mutants exhibited improved thermostability, and the optimal variant, GC20, achieved a melting temperature (Tm) of 76.98 °C, representing a 5.97 °C increase over the wild-type, thereby supporting the practical applicability of SA-MPNN in protein engineering.
Xin-Yue Zhang, Xiang Zheng, Ze-Yuan Dong et al.· Journal of Chemical Informat...· 0 citations
It is concluded that molecular dynamics has an important place in improving the physicality of existing protein structure prediction paradigms, leading to the development of the Subspace Relaxation Operator (SRO).
Colin Baker, Pranav Mahableshwarkar, Ritambhara Singh et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.