Local Geometry Recovers but Cooperative Structure Does Not: Residue-Resolved Limits of All-Atom Reconstruction from Single-Bead Coarse-Grained Disordered Protein Ensembles.
Jul 2026· Journal of Chemical Information and Modeling· Vol 66 16, pp.
10287-10301
· 0 citations· 48 references
Medicine
TL;DR
The benchmark quantifies which atomistic properties can be recovered after projecting a single-bead coarse-grained force field ensemble into all-atom space and identifies improved CG geometric encoding and sequence-conditioned AI priors as the directions for further progress.
Abstract
Intrinsically disordered proteins (IDPs) drive diverse cellular processes through broad conformational ensembles, but experimental characterization of these ensembles is sparse and high-quality training data is scarce, holding back artificial intelligence (AI) approaches to ensemble prediction. Single-bead coarse-grained (CG) force fields such as CALVADOS, parametrized directly against experimental observables, currently provide a more reliable route to disordered ensembles than direct AI prediction. CG sampling lacks atomistic resolution and must be paired with backmapping; the atomistic information recoverable from this two-step process is shaped jointly by the CG representation and the backmapping algorithm, and the interplay between these contributions is not well characterized. We benchmarked CODLAD, a recently published latent-diffusion backmapping pipeline, on 23 Protein Ensemble Database (PED) systems using a dual-input design: the same architecture receives either PED-reference Cα coordinates or independently sampled CALVADOS Cα trajectories. This design separates paired reconstruction error, measurable for PED+CODLAD, from unpaired CG-input-associated ensemble deviations, measurable for CG+CODLAD at the distribution level. Reconstruction from PED conformers achieved 0.56 ± 0.08 Å backbone root-mean-square deviation and preserved local geometry. Reconstruction from CALVADOS-sampled Cα trajectories preserved the Cα framework and bond geometry, while ensemble-level deviations relative to PED were localized to proline backbone geometry and cooperative secondary structure, both consistent with information not carried by an unconstrained single-bead representation. The benchmark quantifies which atomistic properties can be recovered after projecting a CG ensemble into all-atom space and identifies improved CG geometric encoding and sequence-conditioned AI priors as the directions for further progress.
PHASE (Protein Hamiltonians for Sampling of Ensembles), a system-specific framework that converts atomistic conformational ensembles into an explicit and interpretable statistical model, is introduced.
Daniele Angioletti, Marco S. Nobile, Matteo Carli et al.· 0 citations
Transferable coarse-grained (CG) force fields compress chemical space: by aggregating atoms into a reduced set of interaction beads, models such as MARTINI reduce the number of distinguishable compounds by roughly three orders of magnitude, making high-throughput screening of thermodynamic properties tractable across soft matter, with drug--membrane permeability as a well-developed example. The compression is lossy and, so far, one-way: a screen returns a combination of beads, with no established route back to the compounds it stands for. Recovering those compounds--compositional backmapping--is a one-to-many inverse map, distinct from the better-studied conformational problem of rebuilding atomic coordinates from a known mapping. Here we formulate compositional backmapping as conditional graph generation by introducing juniper, a discrete denoising diffusion model over molecular graphs conditioned on the octanol--water partition free energy $\Delta G_{\mathrm{W} \mapsto \mathrm{O}}$, the principal driver of MARTINI bead type assignment and hence a proxy for bead identity. Trained on molecules of up to 9 heavy atoms mapped onto one or two beads, juniper generates molecules that are 93\% valid and 92\% unique for two-bead targets, and whose $\Delta G_{\mathrm{W} \mapsto \mathrm{O}}$ distributions track the target $\Delta G^{\mathrm{CG}}_{\mathrm{W} \mapsto \mathrm{O}}$ linearly ($r^{2} \geq 0.96$), departing only in the hydrophobic and hydrophilic tails. Although the model receives no chemical information beyond a single scalar, the functional groups shift systematically with the imposed free energy, from branched hydrocarbons at the apolar end to amides, imides, and isocyanates at the polar end. A bead combination flagged by a CG screen can therefore be turned into candidate molecules for atomistic study or synthesis.
Luis Itza Vazquez-Salazar, Tristan Bereau· 0 citations
Carbonara is presented, a framework that uses experimental small-angle X-ray scattering (SAXS) data to predict alternative physically plausible protein conformations and provides a route from static structural models of flexible multi-domain proteins and multimeric assemblies to solution-state ensembles.
Pretrained machine-learning interatomic potentials, so-called universal or foundation models offer an appealing starting point for atomistic simulations, but their accuracy for material-specific observables often remains limited without additional reference data (fine-tuning). Here, we systematically quantify how much first-principles data are required to convert universal models into ab initio-accurate material-specific potentials, and ask whether fine-tuning is necessarily preferable to training from scratch. We compare five universal MLIP frameworks, MACE-MP-0, SevenNet-0, GRACE-1L-OAM, MatterSim-v1-5M and ORB-v2, across seven chemically diverse systems incorporating rare and reactive events. Fine-tuning on only 10 AIMD-derived configurations is insufficient for the investigated systems; 200 configurations succeed in favorable cases, but the outcome remains strongly system-dependent. By contrast, 2000 AIMD configurations constitute a robust default, yielding low force and energy errors and reproducing the target material-specific observables. Moderately dense sub-sampling of the AIMD trajectory reduces the required trajectory length tenfold with little loss in model quality. Training from scratch on the same datasets is competitive with, and often slightly more accurate than, naive fine-tuning for MACE and SevenNet, whereas GRACE requires more data. The energy profile for a sulfur-vacancy jump in MoS$_2$ reveals that low trajectory-level errors do not guarantee a correct reaction profile, highlighting the need for observable-level validation. Finally, we show that averaging independently trained models improves predictions in scarce-data regimes at no additional first-principles cost. Together, these results provide practical guidelines for converting limited AIMD reference data into reliable material-specific MLIPs for nanosecond-timescale simulations at near-DFT accuracy.
Gaussian accelerated molecular dynamics (GaMD) enhances conformational sampling by adding a smooth boost potential without requiring predefined collective variables, but an engine-integrated implementation has not been available in GROMACS. Here, we implement total-, dihedral-, and dual-boost GaMD in GROMACS 2025.4, including staged energy-statistics collection, GPU-based bias evaluation and force scaling, restart support, and outputs required for cumulant-based free-energy reweighting. The implementation was evaluated using four benchmark systems spanning conformational free energies, protein folding, and ligand recognition. For alanine dipeptide, a reweighted 100 ns GaMD trajectory recovered the major free-energy basins and rotational barriers in overall agreement with a 1000 ns conventional MD simulation. For chignolin and TC5b, all three independent trajectories for each system sampled native-like folded states from extended conformations within 300 ns and 1 μs, respectively; the best TC5b structure had a minimum backbone RMSD of 0.03 nm from the experimental structure. In the benzene–T4 lysozyme system, two of five independent 500 ns trajectories captured both ligand binding and dissociation, yielding a bound pose with a minimum ligand RMSD of 0.06 nm from the crystal structure. Across all four systems, the boost-potential distributions were approximately Gaussian, and second-order cumulant reweighting resolved the expected conformational and binding free-energy basins. These results demonstrate that GROMACS-GaMD provides a practical, GPU-enabled, collective-variable-free enhanced-sampling framework for biomolecular free-energy calculations, protein folding, and ligand-binding studies.
This work proposes a novel approach that learns hierarchical discrete representations of protein structures using vector quantization, and outperforms state-of-the-art models such as ESMDiff across challenging benchmark datasets, including BPTI MD trajectories and conformational-changing pairs.
Seokjun On, Yujin Jeong, Kanghyeon Kim et al.· Bioinformatics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.