This work introduces mathematical topology metrics that quantify the entanglement complexity of a tertiary protein structure while respecting uncrossability constraints and results indicate that these metrics efficiently encode structural features linked to protein function and provide a more informative description than conventional metrics.
Abstract
Motivation With the rapid development of AI methods that predict protein structures from sequence, understanding the structure-function relation increasingly depends on quantitative structural descriptors that are both biologically meaningful and scalable to large datasets. Here, we introduce mathematical topology metrics that quantify the entanglement complexity of a tertiary protein structure while respecting uncrossability constraints. Results By employing only three such metrics across all protein structures in the Protein Data Bank, we represent the proteome structural space in a continuous three-dimensional space. Distances within this space capture structural similarity and correlate with functional similarity. We find that the mathematical entanglement based landscape of protein structural space diversifies with the evolutionary expansion of protein function across species. Moreover, this continuous representation reproduces CATH classifications with high accuracy for major structural classes. These results indicate that these metrics efficiently encode structural features linked to protein function and provide a more informative description than conventional metrics. Availability Data used in this study are available in the Protein Data Bank. Details of the machine learning model used can be found in https://github.com/roshitac/CATH_Classification-. Contact Banu.Ozkan@asu.edu, Eleni.Panagiotou@asu.edu Supplementary information Supplementary data are available at Journal Name online.
Support Field Neural Representation Learning (SF-NRL), a topology-guided approach that integrates persistent homology(PH), spatial density estimation, and geometric deep learning to infer residue-wise support directly from protein structures, is introduced.
It is argued that, since physics-based simulations and machine learning provide complementary approximations to the underlying probability distribution associated with biomolecular recognition events, and they excel respectively in consistency with free-energy landscapes and state populations and in predictive accuracy, the central challenge for the coming decade will be integrating them into hybrid frameworks that are scalable and transferable.
R. Khalil, Elena Frasnetti, Han Kurt et al.· Journal of Physical Chemistr...· 0 citations
The (un)folding rates of natural proteins determine their native stability and functional homeostasis, making them important targets for protein engineering and design. From a prediction standpoint, the rates have been a long‐standing puzzle. We have known for decades that folding rates empirically correlate with properties of the native three dimensional (3D) structures and that both, folding and unfolding rates, scale with protein size. Whereas such rate correlations are too rough for being of practical use, no significant progress in prediction accuracy has occurred since then, despite many efforts even including machine learning approaches. Here, we retake on this challenge by expanding the simple one‐dimensional free energy surface (1D‐FES) model that originally led to demonstrate the size scaling of both rates, and a curated database with rates for 75 single‐domain proteins. We define the weighted sequence order (WSO) as a novel parameter that allows incorporating structural information into the 1D‐FES model explicitly. Via the WSO, we examine the role of global structural properties such as fold topology and core packing in defining the (un)folding rates within the context of a physics‐based model of protein folding. After introducing fold topology and packing at a coarse‐grained level, the model uses three floating parameters to predict the folding and unfolding rates within 6.5‐ and 10‐fold, respectively, resulting in ±6.5 kJ/mol accuracy in native stability, equivalent to the typical perturbation induced by one single‐point mutation. The net improvement over the 2‐parameter size‐only prediction is of 2.5‐fold. These new rate predictions are significantly closer to the threshold of usefulness for engineering and design. More importantly, this WSO‐modified 1D‐FES model can now directly accommodate atomistic, high‐resolution, force‐fields to further optimize the rate predictions, and/or to use rate information as a testbed for force‐field refinement. Finally, the WSO‐1D‐FES model could also serve as foundation for developing more complex models capable of dealing with multi‐domain proteins as well as with the evolutionary information cryptically encoded in natural protein sequences.
Mohammad Abdulqader, Victor Muñoz· Protein Science· 0 citations
The answer lies in geometry: proteins with denser cores, larger size, and higher-order oligomeric assembly tolerate mutations more readily, occupy larger structural families, and support more versatile biological roles, reveals that protein size, shape, and self-assembly, not just sequence, are fundamental drivers of evolvability.
By smoothing the Evoformer's weight tensors with a Gaussian convolution and scaling the result, it is shown that the trained model produces physically structured conformational landscapes, appearing to encode structural constraints that extend beyond what unperturbed inference reveals.
An improved force field is developed, derived from its parent, Amber ff24EXP-GA, and its evaluation against Amber ff14SB and other contemporary force fields, such as CHARMM36m, in capturing the empirically determined conformational properties of unfolded systems: short peptides that serve as model systems for IDPs, and longer unfolded proteins.
Athul Suresh, B. Urbanc· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.