Skip to content

AquaGen: Scaling generative models to molecular dynamics precision on thousands of atoms

Jul 2026 · arXiv.org · Vol abs/2607.03513 · 1 citation · 52 references
Physics Computer Science

TL;DR

AquaGen is presented, the first all-atom, explicit solvent, periodic-boundary-condition-aware generative model that produces molecular configurations from the Boltzmann distribution at a fraction of the cost of molecular dynamics (MD) and demonstrates the utility of high-resolution ensemble generation for free energy estimation.

Abstract

We present AquaGen, the first all-atom, explicit solvent, periodic-boundary-condition-aware generative model that produces molecular configurations from the Boltzmann distribution at a fraction of the cost of molecular dynamics (MD). This is in contrast with existing generative models that remove degrees of freedom by operating on coarse-grained, vacuum, or implicit solvent systems. Operating at this resolution allows for post-processing through force field energy evaluations and MD simulations, and enables the prediction of relevant properties in a gray-box manner (as ensemble averages of potential energy evaluations over generated samples). We demonstrate the utility of this paradigm on absolute hydration free energy (AHFE), producing estimates 4-10x faster and with comparable accuracy to standard GPU-based MD. By generating uncorrelated samples from alchemical Boltzmann distributions, we create more accurate, interpretable, and refinable ensemble predictions with calibrated uncertainty estimates, unlike regression methods which are entirely black-box predictors. Our approach also yields predictable benefits from increasing train- and test-time compute, realized by scaling model size and generating more samples, respectively. We believe that this approach demonstrates the utility of high-resolution ensemble generation for free energy estimation, with future potential to replace MD in tasks such as the prediction of lipophilicity, membrane permeability, or absolute binding free energy (ABFE) -- whose grounding and interpretability may be critical for the development of new drugs and materials.

View source

Similar papers

Jul 2026

Implicit Machine Learning Force Fields Accelerate Molecular Dynamics Simulations

This work introduces implicit machine learning force fields, which replace explicit stacks of neural network layers with self-consistent fixed-point equations, and demonstrates this across three major classes of graph neural networks: invariant, equivariant Cartesian tensor, and SO(3)-equivariant spherical-tensor architectures.

J. Maeß, Leon Werner, J. Frank et al. · 0 citations
#machine learning Preprint Sep 2026

Generative Nested Sampling of Atomistic Thermodynamic Landscapes

Nested sampling (NS) resolves the thermodynamics of an atomistic system from a single simulation, but its practical reach is limited by the Markov-chain updates needed to decorrelate walkers within each likelihood-constrained ensemble. Flow-based NS has removed this bottleneck for gravitational-wave (GW) inference, yet its transfer to atomistic systems is not merely a change of application. Comparing a GW150914-like binary-black-hole likelihood with an eight-particle two-dimensional Lennard-Jones (LJ) system of comparable dimensionality, we show that the two landscapes differ fundamentally: atomistic multimodality is discrete and combinatorial, generated by particle permutations separated by hard collision walls, and its coordinate coupling is dense and collective, whereas the GW posterior exhibits smooth degeneracies and localized parameter coupling. Guided by this diagnosis, we introduce NS-Flows: a single conditional normalizing flow, conditioned on the NS energy bound and trained on a sliding window of recent live sets, that replaces MCMC by direct parallel draws corrected by importance-weighted rejection resampling. Live sets supply data self-consistently, allowing flow training without structured priors or a pre-existing dataset. For LJ disks in PBC, the algorithm reduces energy evaluations by over two orders of magnitude and wall-clock time by roughly one third, an advantage that becomes increasingly favorable as the cost of the potential grows. The flow's generation efficiency further acts as a physical diagnostic: it varies non-monotonically along the annealing trajectory, is lowest in the dense disordered regime, and is quantitatively captured by the constrained ensemble's internal mode complexity together with target drift across the training window, identifying liquid-like ensembles, rather than prior-target separation, as the hard case for current flow architectures.

A. Coretti, N. Unglert, Sebastian Falkner et al. · 0 citations
#protein folding Sep 2026

Convergence Mechanisms of Generative Models in Molecular Conformational Sampling

Characterizing equilibrium conformational ensembles with deep generative models requires understanding whether a model reproduces a target distribution and how it reaches that distribution. Here, we compare two generative routes to molecular conformational sampling, stochastic relaxation and deterministic transport, using denoising diffusion probabilistic models and rectified-flow models across systems of increasing complexity: a multimodal two-dimensional potential, the folded miniprotein Trp-cage, and a high-dimensional dihedral representation of an intrinsically disordered protein. We show that these paradigms differ in end point fidelity and in how distributional error is resolved during sampling. Diffusion models converge through pronounced late-stage stochastic relaxation and robustly recover the configurational breadth across neural architectures. Rectified flow approaches the target distribution through deterministic transport and therefore depends more strongly on architectural expressivity, particularly in heterogeneous, high-dimensional landscapes. Entropy and moment-evolution analyses further show that diffusion more reliably restores the ensemble location and fluctuation structure, whereas rectified flow requires Transformer-level feature mixing to represent transport geometry accurately. These results establish the convergence mechanism as a practical design principle for molecular generative sampling, clarifying when stochastic diffusion provides robustness and when deterministic transport requires higher representational capacity.

Nagesh B E, Jagannath Mondal · 0 citations
Book Open access Jul 2026

Moore’s Law for Molecular Dynamics Simulations: An Initial Study

High-performance computing (HPC) is essential to modern molecular dynamics (MD) simulations. While computing capabilities continue to grow, it remains unclear how these advances translate into everyday scientific practice. Additional resources may be used to reduce time-to-solution, improve model fidelity, simulate larger systems, extend simulated time scales, or increase the number of independent runs. Understanding how users allocate increased computational capability is important for future system design. To investigate how MD usage has evolved over time, we analyzed over 400 publications reporting all-atom MD simulations using a large language model (LLM)-based extraction workflow. Thirty of these publications were manually reviewed to develop and validate an expert-guided extraction rubric. Using this rubric, we evaluated several LLM configurations for extracting system size, simulation duration, and number of independent runs. Extraction quality improved across successive model generations, from GPT-4o to GPT-5 and then to GPT-5.4. Within the GPT-5.4 family, accuracy decreased as model size was reduced from GPT-5.4 to GPT-5.4-mini and then to GPT-5.4-nano. However, increasing reasoning effort from medium to high made GPT-5.4-nano performance comparable to GPT-5.4 while reducing cost by nearly an order of magnitude. Based on this result, GPT-5.4-nano with high reasoning effort was selected for extraction across the full dataset. Our initial analysis suggests that MD system sizes and the number of independent runs have not grown substantially over the last two decades, whereas simulation duration shows a clearer, approximately exponential increase. The fitted trend suggests that maximum MD simulation duration doubled approximately every 2.2 years between 2005 and 2025. These results demonstrate the feasibility of AI-assisted meta-analysis for understanding historical trends in scientific computing workloads.

Alexey N. Simakov, Nikolay A. Simakov · 0 citations
Preprint Aug 2026

Universal Machine-learning Molecular Dynamics at the Speed of Empirical Potentials

DPA4C is introduced, an equivariant potential whose architecture and compressed CUDA operators are co-designed under deployment constraints to pursue accuracy and efficiency together and brings quantum-trained universal accuracy into a regime of speed and system size previously associated with empirical potentials.

Tian Li, Jianming Xue, Linfeng Zhang et al. · 0 citations
Review Open access Jul 2026

Predicting Biomolecular Interactions in the Next Decade: Physics-Based Methods Meet AI-Driven Approaches.

It is argued that, since physics-based simulations and machine learning provide complementary approximations to the underlying probability distribution associated with biomolecular recognition events, and they excel respectively in consistency with free-energy landscapes and state populations and in predictive accuracy, the central challenge for the coming decade will be integrating them into hybrid frameworks that are scalable and transferable.

R. Khalil, Elena Frasnetti, Han Kurt et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.