Skip to content
Open access

Carbonara: a SAXS-guided seeding framework for exploring protein solution-state dynamics

Jul 2026 · bioRxiv · 0 citations · 72 references
Biology

TL;DR

Carbonara is presented, a framework that uses experimental small-angle X-ray scattering (SAXS) data to predict alternative physically plausible protein conformations and provides a route from static structural models of flexible multi-domain proteins and multimeric assemblies to solution-state ensembles.

Abstract

Proteins in solution often populate conformational ensembles that differ from the static states captured by crystallography or AI-based structure prediction. Conventional molecular dynamics (MD) simulations often fail to cross the energy barriers separating these states on accessible timescales, and statistical reweighting cannot recover conformations never sampled. Here we present Carbonara, a framework that uses experimental small-angle X-ray scattering (SAXS) data to predict alternative physically plausible protein conformations. Carbonara builds on Wiggle, a standalone Cα-based SAXS forward model validated against explicit-solvent calculations and experimental benchmarks. Using two case studies, an AI-predicted multi-domain helicase (SMAR-CAL1) and a crystallographic antibody fragment (ChiLob7/4 IgG2), we show how seeding MD simulations from Carbonara conformations enables efficient exploration of solution-state conformational landscapes. In both cases, MD ensembles initiated from available models either fail to match the SAXS data or do so only after discarding nearly all sampled conformations, whereas Carbonara-seeded ensembles reach agreement while retaining the majority of conformations. Our modelling framework provides a route from static structural models of flexible multi-domain proteins and multimeric assemblies to solution-state ensembles.

Read PDF

Similar papers

Open access Aug 2026

Pi-Ensemble: Sequence-guided generation of interpolated protein conformational ensembles

Pi-Ensemble (Predicting Interpolated Ensemble), a sequence-guided framework for generating protein conformational ensembles interpolating between two structural anchor states, provides an extensible framework for studying protein flexibility, guiding adaptive sampling, and accelerating mechanistic investigations of protein function.

Hassan Nadeem, D. Kleiman, Yu-Ming Zhou et al. · 0 citations
Open access Aug 2026

Benchmarking AI-generated structural ensembles of membrane proteins against physics-based modelling

It is demonstrated that BioEmu can generate plausible conformational ensembles for relatively large, six-and seven-pass membrane proteins, sampling rare states at a fraction of the computational cost of conventional MD simulations, suggesting that AI-based ensemble generation could provide an accessible approach for exploring membrane protein dynamics and complement conventional molecular modelling approaches.

B. Clifton, Adam G Grieve, Robin A. Corey · 0 citations
Open access

Experimental data-guided parameterization and validation of an AMBER protein force field

An improved force field is developed, derived from its parent, Amber ff24EXP-GA, and its evaluation against Amber ff14SB and other contemporary force fields, such as CHARMM36m, in capturing the empirically determined conformational properties of unfolded systems: short peptides that serve as model systems for IDPs, and longer unfolded proteins.

Athul Suresh, B. Urbanc · 0 citations
Open access Sep 2026

A Simulation-Free Topological Basis for Building Compact Koopman Models of Protein Folding

Unravelling protein-folding mechanisms and kinetics is a key challenge to biochemical science. The variational approach for Markov processes (VAMP) is a powerful tool to build Markov Models that capture key kinetic and structural information despite the conformational complexity and long time scales associated with protein folding. However, VAMP-based Markov models use data from exhaustive molecular dynamics (MD) simulations to construct an underlying basis set describing the “coarse-grained” kinetics; the same MD data can be used to predict transition probabilities between partitioned configuration space, enabling the extraction of folding time scales and mechanism. Here, we propose an alternative strategy for Markov model construction that does not rely on extensive, computationally demanding MD data for configuration space partitioning. Specifically, we show that graph-driven sampling (GDS) can generate a complete “landscape” of intermediate contact-maps linking unfolded and folded protein conformations; importantly, extensive MD simulations are not required in GDS. When combined with a physically intuitive shortest-contact-hop metric to discriminate different intermediate states, GDS mapping of protein-folding configuration space generates a reliable “structurally aware” partitioning for VAMP model construction. To demonstrate this strategy, we show that a GDS-constructed Markov model variationally improves folding time scale estimates for all-atom models of the WW domain protein─and has the additional advantage of easily resolving kinetic traps in the folding landscape that have proven challenging to confirm otherwise. Together, the combination of GDS and VAMP opens a new route toward rapid characterization of protein-folding intermediates and kinetics traps to help address frontier challenges such as protein misfolding, aggregation, and protein design.

Ziad Fakhoury, G. Sosso, S. Habershon · 0 citations
Open access Aug 2026

Inferring protein ensembles directly from NOESY spectra

A quantitative scoring framework for comparing experimental and back-calculated observables is introduced and combined with regularized ensemble selection and Monte Carlo simulated annealing to provide direct inference of protein ensembles within a flexible ensemble-selection architecture incorporating multiple classes of NMR observables.

Murray Coles · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.