This work reviews the computational and experimental approaches that disentangle folding and function at scale, revealing a dark energy component and providing new insights into how biological information flows from sequence to structure to function and back to sequence.
Abstract
The evolutionary fate of proteins is driven by both folding stability and biological function, dual constraints that often conflict, creating frustration and imposing functional costs beyond stability. These costs can be captured by a"dark energy": the difference between the evolutionary energy of protein sequences and their physical folding energy. Recent advances in deep mutational scanning, protein language models, and inverse-folding models have enabled the quantification of dark energy across the protein universe. We review the computational and experimental approaches that disentangle folding and function at scale, revealing a dark energy component and providing new insights into how biological information flows from sequence to structure to function and back to sequence.
Together, these insights position conformational dynamics at the center of understanding and engineering the evolutionary logic of protein function, opening the door to study how proteins are tuned to operate under the nonequilibrium conditions of living cells.
Sixto M. Herrera, Elías Manríquez-Benítez, Exequiel Medina· Current Opinion in Structura...· 0 citations
It is concluded that the early emergence of the Rossmann fold reflects the chemical and physical constraints of protein folding, explaining both its profound antiquity and sustained longevity.
Koh Seya, Tatsuya Corlett, Hamza Giaffar et al.· bioRxiv· 0 citations
The (un)folding rates of natural proteins determine their native stability and functional homeostasis, making them important targets for protein engineering and design. From a prediction standpoint, the rates have been a long‐standing puzzle. We have known for decades that folding rates empirically correlate with properties of the native three dimensional (3D) structures and that both, folding and unfolding rates, scale with protein size. Whereas such rate correlations are too rough for being of practical use, no significant progress in prediction accuracy has occurred since then, despite many efforts even including machine learning approaches. Here, we retake on this challenge by expanding the simple one‐dimensional free energy surface (1D‐FES) model that originally led to demonstrate the size scaling of both rates, and a curated database with rates for 75 single‐domain proteins. We define the weighted sequence order (WSO) as a novel parameter that allows incorporating structural information into the 1D‐FES model explicitly. Via the WSO, we examine the role of global structural properties such as fold topology and core packing in defining the (un)folding rates within the context of a physics‐based model of protein folding. After introducing fold topology and packing at a coarse‐grained level, the model uses three floating parameters to predict the folding and unfolding rates within 6.5‐ and 10‐fold, respectively, resulting in ±6.5 kJ/mol accuracy in native stability, equivalent to the typical perturbation induced by one single‐point mutation. The net improvement over the 2‐parameter size‐only prediction is of 2.5‐fold. These new rate predictions are significantly closer to the threshold of usefulness for engineering and design. More importantly, this WSO‐modified 1D‐FES model can now directly accommodate atomistic, high‐resolution, force‐fields to further optimize the rate predictions, and/or to use rate information as a testbed for force‐field refinement. Finally, the WSO‐1D‐FES model could also serve as foundation for developing more complex models capable of dealing with multi‐domain proteins as well as with the evolutionary information cryptically encoded in natural protein sequences.
Mohammad Abdulqader, Victor Muñoz· Protein Science· 0 citations
Foldable proteins exhibit funnel-like energy landscapes, but their evolution under selection for function remains unclear. We study this in a lattice protein model, defining fitness as the equilibrium probability of a fixed active-site motif and using multicanonical sequence sampling. At intermediate temperature, rare high-fitness sequences show funnel-like energy and double-well free-energy landscapes. At very low temperature, abundant high-fitness sequences retain glass-like landscapes. Thus, thermal fluctuations turn a local functional requirement into a global structural constraint.
It is shown that short nucleic acids containing Gquadruplex (G4) structure can also catalyze protein folding and uncovers a previously underappreciated role for nucleic acid in proteostasis and offers a new strategy for studying nucleic acid structure-function relationship at residue level.
It is argued that incorporating frustration into computational and experimental strategies will be essential to move beyond purely stability-driven approaches toward the rational engineering of functional proteins.
Franco L. Simonetti, Eli J. Draizen, Rocío Espada et al.· Biochimica et Biophysica Act...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.