Skip to content
Open access

Blending physics-based and inverse folding models to disentangle variant effects on stability and function

Aug 2026 · bioRxiv · 0 citations · 79 references
Biology

TL;DR

It is shown that blending IF models with a physics-based coarse-grained potential improves global correlation with experimental ΔΔG and, crucially, reduces IF model bias at functional sites, and is found that disease gain-of-function variants show a distinct functional signature from loss-of-function variants.

Abstract

Protein sequences are constrained not only by the need to fold into stable structures, but also by specific functional requirements imposed by natural selection. Yet predictions of how amino-acid changes affect proteins typically collapse these constraints into a single scalar score. Quantitatively separating these effects at scale remains an open challenge, with direct relevance spanning protein design to understanding the molecular mechanisms of disease. Inverse-folding (IF) models have emerged as fast, unsupervised predictors of folding energy changes (ΔΔG), but because they learn statistical correspondences between structure and sequence, they can conflate conservation driven by function with conservation driven by stability. Here, we show that blending IF models with a physics-based coarse-grained potential improves global correlation with experimental ΔΔG and, crucially, reduces IF model bias at functional sites. Applying the best-performing blend together with an evolutionary language model, we decompose each variant’s evolutionary cost into folding energy and dark energy, the latter capturing functional constraints beyond folding stability. With this decomposition, and without the need for supervision, we find that disease gain-of-function variants show a distinct functional signature from loss-of-function variants. In particular, we identify oncogenic drivers as largely preserving stability while exhibiting high dark energy, as opposed to tumor suppressors which are predominantly destabilized, paving the way to a mechanistic understanding of driver mutations in cancer. Together, these results provide a scalable framework for accurate ΔΔG prediction and mechanistic disentanglement of variant effects.

Read PDF

Similar papers

Open access Jul 2026

Physics‐based prediction of protein folding and unfolding rates: Examining the roles of fold topology and core packing

The (un)folding rates of natural proteins determine their native stability and functional homeostasis, making them important targets for protein engineering and design. From a prediction standpoint, the rates have been a long‐standing puzzle. We have known for decades that folding rates empirically correlate with properties of the native three dimensional (3D) structures and that both, folding and unfolding rates, scale with protein size. Whereas such rate correlations are too rough for being of practical use, no significant progress in prediction accuracy has occurred since then, despite many efforts even including machine learning approaches. Here, we retake on this challenge by expanding the simple one‐dimensional free energy surface (1D‐FES) model that originally led to demonstrate the size scaling of both rates, and a curated database with rates for 75 single‐domain proteins. We define the weighted sequence order (WSO) as a novel parameter that allows incorporating structural information into the 1D‐FES model explicitly. Via the WSO, we examine the role of global structural properties such as fold topology and core packing in defining the (un)folding rates within the context of a physics‐based model of protein folding. After introducing fold topology and packing at a coarse‐grained level, the model uses three floating parameters to predict the folding and unfolding rates within 6.5‐ and 10‐fold, respectively, resulting in ±6.5 kJ/mol accuracy in native stability, equivalent to the typical perturbation induced by one single‐point mutation. The net improvement over the 2‐parameter size‐only prediction is of 2.5‐fold. These new rate predictions are significantly closer to the threshold of usefulness for engineering and design. More importantly, this WSO‐modified 1D‐FES model can now directly accommodate atomistic, high‐resolution, force‐fields to further optimize the rate predictions, and/or to use rate information as a testbed for force‐field refinement. Finally, the WSO‐1D‐FES model could also serve as foundation for developing more complex models capable of dealing with multi‐domain proteins as well as with the evolutionary information cryptically encoded in natural protein sequences.

Mohammad Abdulqader, Victor Muñoz · 0 citations
#artificial intelligence Preprint Sep 2026

SymFold: Synergizing Evolutionary and Structural Priors for Accurate Protein Inverse Folding

Protein inverse folding aims to recover amino acid sequences for a given 3D protein structure, underpinning broad applications such as enzyme engineering and drug discovery.Current methods often follow a serial pipeline, in which a structure encoder predicts a coarse sequence, which is then refined by protein language models (PLMs). However, because PLMs only perform post-hoc sequence edits, the refinement is bounded by the quality of upstream predictions.Thanks to recent multimodal protein language models (MPLMs), we could directly encode structure to generate sequences with pretrained structural knowledge, but we observe that they are not effective for inverse folding. Therefore, we introduce a symmetric dual-path architecture that both leverages PLMs for pretrained sequence evolution knowledge and MPLMs for pretrained structural knowledge to iteratively guide protein sequence generation.Through extensive experiments across standard protein inverse folding benchmarks, our method achieves state-of-the-art performance, surpassing prior approaches, and ablation studies validate the rationale of our symmetric design, revealing a promising direction for the community.

Han-Dong Wang, Jiaxin Qi, Baisheng Lai et al. · 0 citations
Open access Sep 2026

A Simulation-Free Topological Basis for Building Compact Koopman Models of Protein Folding

Unravelling protein-folding mechanisms and kinetics is a key challenge to biochemical science. The variational approach for Markov processes (VAMP) is a powerful tool to build Markov Models that capture key kinetic and structural information despite the conformational complexity and long time scales associated with protein folding. However, VAMP-based Markov models use data from exhaustive molecular dynamics (MD) simulations to construct an underlying basis set describing the “coarse-grained” kinetics; the same MD data can be used to predict transition probabilities between partitioned configuration space, enabling the extraction of folding time scales and mechanism. Here, we propose an alternative strategy for Markov model construction that does not rely on extensive, computationally demanding MD data for configuration space partitioning. Specifically, we show that graph-driven sampling (GDS) can generate a complete “landscape” of intermediate contact-maps linking unfolded and folded protein conformations; importantly, extensive MD simulations are not required in GDS. When combined with a physically intuitive shortest-contact-hop metric to discriminate different intermediate states, GDS mapping of protein-folding configuration space generates a reliable “structurally aware” partitioning for VAMP model construction. To demonstrate this strategy, we show that a GDS-constructed Markov model variationally improves folding time scale estimates for all-atom models of the WW domain protein─and has the additional advantage of easily resolving kinetic traps in the folding landscape that have proven challenging to confirm otherwise. Together, the combination of GDS and VAMP opens a new route toward rapid characterization of protein-folding intermediates and kinetics traps to help address frontier challenges such as protein misfolding, aggregation, and protein design.

Ziad Fakhoury, G. Sosso, S. Habershon · 0 citations
Book Open access Aug 2026

From Structure to Function: Preference Alignment for Function-aware Protein Inverse Folding

A function-aware preference alignment framework that improves functional preservation by fine-tuning models to favor function-preserving sequences over function-disrupting alternatives, avoiding the need for explicit function optimization.

Nilufer Tamatgar, Soobin Park, Yinghua Yao et al. · 0 citations
Preprint Sep 2026

Predicting directional flexibility in proteins

Predicting protein dynamics is a long-standing problem in computational structural biology. Often, protein function critically depends on local directed motions, such as hinge movements, catalytic loop rearrangements and domain reorientations, which can be characterized by directional flexibility and correlated structural motions of the protein backbone. While Molecular Dynamics (MD) simulations provide an established but often prohibitively expensive approach, recent deep generative models aim to reduce this cost by directly predicting conformational ensembles, emulating MD. However, due to their large size and the need to generate several states until the derived dynamical properties converge, these models remain expensive. In this work, we propose BackFlip-2: a fast SE(3)-equivariant graph neural network trained to directly predict dynamical descriptors, such as directional backbone flexibility and pairwise dynamic correlations, from an equilibrium structure. In a series of experiments, we show that our model matches the accuracy of substantially larger ensemble generation models while being orders of magnitude faster, and demonstrate that the proposed equivariant architecture is especially well-suited for capturing anisotropic motions in proteins. BackFlip-2 model weights, training and inference code are available at https://github.com/graeter-group/backflip.

Vsevolod Viliuga, Leif Seute, Matteo Tadiello et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.