SPINDLE, an ensemble of deep neural networks trained on a large synthetic set of NMR relaxation data predicts both fast and slow timescale dynamics parameters from a single set of three relaxation datasets using the ensemble for error estimation, and demonstrates a strong correlation to ground truth dynamics parameters on synthetic benchmarks.
Abstract
A protein’s function is derived from its three-dimensional structure and the motions of the atoms about that structure. The detailed characterization of both macromolecular structure and dynamics provides an opportunity for understanding enzyme catalysis, ligand binding, and allostery, along with providing insights into how the function changes upon mutation or post-translational modification. Among the various methods for characterizing biomolecular motions, nuclear magnetic resonance (NMR) spin relaxation methods are a standard for determining nanosecond global tumbling times along with the amplitude and timescale of faster local motions. Within the model-free formalism, various mathematical models are used to extract dynamic parameters. Unfortunately, as the number of fitted parameters increases within these models, they become mathematically underdetermined for standard NMR relaxation data collected at a single magnetic field, necessitating multi-field datasets. Here, we present SPINDLE, an ensemble of deep neural networks trained on a large synthetic set of NMR relaxation data. Unlike traditional least-squares fitting, SPINDLE predicts both fast and slow timescale dynamics parameters from a single set (i.e., collected at a single magnetic field) of three relaxation datasets using the ensemble for error estimation. We demonstrate a strong correlation to ground truth dynamics parameters on synthetic benchmarks, with more precision than traditional fitting techniques, and precisely reproduce experimental dynamics parameters for ∼50 proteins with relaxation data in the Biological Magnetic Resonance Data Bank. We also leverage the architecture of the deep neural network to show how the model emphasizes rigid residues for the prediction of global correlation times. This strategy may be useful in the future for elucidating correlated networks of dynamic residues from multiple relaxation datasets.
A quantitative scoring framework for comparing experimental and back-calculated observables is introduced and combined with regularized ensemble selection and Monte Carlo simulated annealing to provide direct inference of protein ensembles within a flexible ensemble-selection architecture incorporating multiple classes of NMR observables.
Nuclear magnetic resonance (NMR) spectroscopy yields rich residue-level information on biomolecular dynamics and chemical environments, two frontiers for quantitative predictive methods in biochemistry. Decades of data are publicly archived in the Biological Magnetic Resonance Data Bank (BMRB)1, yet in practice, this information remains difficult to access and interpret at scale and within computational workflows. Here we present makeshift, an open-source Python package for accessing, curating, and analyzing NMR datasets. Users can readily retrieve and parse BMRB entries and perform essential analyses such as chemical shift re-referencing, secondary structure propensity prediction, and interpretation of relaxation datasets for dynamics. We re-implemented several widely-used NMR data calculations which were not open-source or available in Python and validated our implementations against the original implementations. By integrating data access, processing, and analysis into a single Python interface, makeshift lowers the barrier for reproducible, scalable analysis and machine learning applications using biomolecular NMR data.
Gina El Nesr, Hannah K. Wayment-Steele· bioRxiv· 0 citations
It is argued that, since physics-based simulations and machine learning provide complementary approximations to the underlying probability distribution associated with biomolecular recognition events, and they excel respectively in consistency with free-energy landscapes and state populations and in predictive accuracy, the central challenge for the coming decade will be integrating them into hybrid frameworks that are scalable and transferable.
R. Khalil, Elena Frasnetti, Han Kurt et al.· Journal of Physical Chemistr...· 0 citations
P pHaseMD4AI is presented, a molecular dynamics dataset that combines a globally equilibrated peptide branch with a protein-scale constant-pH molecular dynamics (CpHMD) branch spanning hundreds of soluble proteins and provides a resource for developing and benchmarking molecular machine learning methods.
Tiefeng Song, Yi-Xin Guo, Jiahao He et al.· bioRxiv· 0 citations
Pi-Ensemble (Predicting Interpolated Ensemble), a sequence-guided framework for generating protein conformational ensembles interpolating between two structural anchor states, provides an extensible framework for studying protein flexibility, guiding adaptive sampling, and accelerating mechanistic investigations of protein function.
Hassan Nadeem, D. Kleiman, Yu-Ming Zhou et al.· bioRxiv· 0 citations
Proteins are intrinsically dynamic molecules that continuously explore conformational ensembles to execute biological functions. Conventional structural biology methods rely on in vitro reconstitution of purified components and therefore capture predominantly static snapshots, often overlooking the regulatory roles of the cellular microenvironment, such as molecular crowding, weak interaction networks, and post-translational modifications. This limitation has driven an urgent need to transition from in vitro reconstruction to in vivo characterization within living cells. Nuclear magnetic resonance (NMR) spectroscopy provides atomic-resolution insights into structure and motions spanning multiple timescales, yet its application is constrained by molecular weight limits, isotopic labeling requirements, and inherently low throughput. Cross-linking mass spectrometry (XL-MS) complements NMR by delivering sparse but long-range spatial restraints without an upper molecular weight limit. The integration of NMR and XL-MS establishes a powerful synergistic framework that bridges atomic-resolution local structures and large-scale interaction topologies, thereby enabling comprehensive characterization of protein dynamic conformations and interaction networks in native environments. Here, we review how this integrative strategy advances the understanding of intrinsically disordered proteins, multi-domain proteins, and dynamic protein-protein interaction networks in native cellular environments. We further discuss emerging technological frontiers, including hyperpolarized NMR, photo-cross-linking, organelle-resolved analysis, and artificial intelligence-guided integrative modeling, which together promise to transform our ability to resolve the true functional states of proteins inside cells.
Zhou Gong, Qun Zhao, Min Sun et al.· Magnetic Resonance Letters· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.