Skip to content

Enhanced Sampling in the Age of Machine Learning: Algorithms and Applications

Sep 2025 · Chemical Reviews · Vol 126, pp. 671 - 713 · 58 citations · 327 references
Medicine Physics

TL;DR

A comprehensive overview of how enhanced sampling methods are reshaping the field, with a particular focus on the data-driven construction of collective variables, is provided.

Abstract

Molecular dynamics simulations hold great promise for providing insight into the microscopic behavior of complex molecular systems. However, their effectiveness is often constrained by long timescales associated with rare events. Enhanced sampling methods have been developed to address these challenges, and recent years have seen a growing integration with machine learning techniques. This Review provides a comprehensive overview of how they are reshaping the field, with a particular focus on the data-driven construction of collective variables. Furthermore, these techniques have also improved biasing schemes and unlocked novel strategies via reinforcement learning and generative approaches. In addition to methodological advances, we highlight applications spanning different areas, such as biomolecular processes, ligand binding, catalytic reactions, and phase transitions. We conclude by outlining future directions aimed at enabling more automated strategies for rare-event sampling.

Read PDF

Similar papers

Review Aug 2026

Leveraging generative models to assist Monte Carlo sampling

Sampling high-dimensional probability distributions is a central task in scientific computing, with applications ranging from Bayesian inference to statistical physics and molecular simulation. Despite decades of methodological developments, two major challenges remain: scaling to high dimensions and efficiently exploring multimodal distributions characterized by metastable states. Classical approaches such as Markov chain Monte Carlo, tempering methods, or enhanced sampling based on collective variables have achieved major successes, but they also face intrinsic limitations. This tutorial review explores a new paradigm that has recently emerged at the interface of machine learning and computational statistical physics: the use of generative models as tools for sampling. In this context, models such as normalizing flows and diffusion models are not used in their traditional data-driven setting, but rather as flexible probabilistic models that can assist the sampling of distributions known only up to a normalization constant. This manuscript reviews the early development of this rapidly evolving field and discusses several methodological directions, including exact samplers based on generative models and strategies to train such models in the absence of data. While an exhaustive survey of the literature is not attempted, we present a selection of key ideas and methods, along with a discussion of their strengths and limitations. The review is intended to be an accessible tutorial for both physics and machine learning audiences, and it aims to provide a starting point for researchers interested in exploring this exciting area of research.

Marylou Gabrié · 0 citations
Open access Aug 2026

AdaptivePy: a unified Python framework for adaptive sampling in molecular dynamics

Adaptive sampling accelerates the exploration of conformational space in molecular dynamics (MD) simulations by repeatedly analyzing the accumulated trajectories and seeding a new round of simulations from informative configurations. A growing collection of adaptive sampling policies has been proposed, each built around a particular notion of what makes a configuration informative, yet these methods are scattered across separate and often incompatible implementations, which complicates their systematic comparison and their combined use in meta adaptive sampling schemes. Here, we present AdaptivePy, a compact and extensible Python framework that implements nine seed-selection policies behind a single configuration-driven interface, spanning simple population-based baselines, several established machine-learning and geometry-based methods, and two ensemble or meta sampling policies introduced in this work. We show that the shared implementation reproduces the characteristic selection behavior of each policy on a series of analytic benchmark landscapes. We also introduce a new adaptive sampling scheme that employs TS-DAR, a deep learning framework originally designed to identify transition states, into an acquisition criterion that drives the discovery of an entire multi-basin landscape starting from a single basin. We further demonstrate that the common interface enables meta adaptive sampling policies, which aggregate the rankings of several policies into a single set of seeds. AdaptivePy thereby provides a unified testbed for the adoption, benchmarking, and continued development of adaptive sampling methods for biomolecular MD simulations.

Hassan Nadeem, D. Kleiman, Diwakar Shukla · 0 citations
Open access Aug 2026

Multiscale and Multi‐Timestep Switching of Multiple Machine Learning Force Fields for Artificial Intelligence‐Driven Materials Simulations

Molecular dynamics (MD) is essential for investigating atomic‐scale processes in materials and molecular systems, but the cost of high‐accuracy machine learning force field simulations still limits accessible system sizes and timescales. Here, we propose a practical model‐switching strategy for Deep Potential (DP)‐based MD simulations that alternates between independently trained DP models with different cutoff radii: a standard 6 Å model for higher accuracy and a reduced‐cutoff 4 Å model for faster inference. The method was implemented in LAMMPS/DeePMD and evaluated using solid‐phase anatase TiO 2 and liquid‐phase polyethylene glycol (PEG). For anatase TiO 2 , the 1:3 4–6 Å switching scheme preserved radial distribution function (RDF) correlations of 0.996 or higher relative to the 6 Å baseline while achieving a 1.24‐fold speedup. For PEG, the switching scheme maintained RDF correlations of 0.996 or higher with a 1.18‐fold speedup. Additional optimization using network‐size reduction and mixed‐precision inference achieved a 2.53‐fold speedup with RDF correlations of 0.975–0.988. Constant particle‐number, pressure, and temperature (NPT) simulations remained stable, whereas constant particle‐number, volume, and energy (NVE) simulations revealed system‐dependent energy‐drift behavior, particularly for aggressively optimized models. These results demonstrate that DP model switching provides a simple and practical route for accelerating structural MD simulations while highlighting the need for validation when strict energy conservation is required.

R. Kanda, Megumu Yamazaki, Yuta Yoshimoto et al. · 0 citations
Preprint Aug 2026

Adaptive Inference and Convergence of Free Energy Landscapes Using Non-parametric Bayesian Enhanced Sampling

Enhanced sampling techniques are central to the study of statistically rare events in the computer modeling and simulation of molecular phenomena. In this work, we report the development and integration of a Gaussian process model, adaptive, uncertainty-driven sampling scheme for enhanced sampling. The framework trains a Gaussian process model on an iteratively improving estimate of the free energy landscape. Regions of high uncertainty within the reaction phase space are increasingly sampled, and the uncertainty is computed on-the-fly during free energy reconstruction, serving as a convergence metric. This approach provides a generalizable strategy that can be extended to complex molecular processes.

D. Kamp, Sinai Lee, R. Phung et al. · 0 citations
Open access

Simulation intelligence for rare events in biomolecular processes

Molecular dynamics (MD) simulations are powerful tools for investigating the dynamic behavior of complex biomolecular systems at atomistic resolution. However, the size and complexity of the configuration space pose substantial challenges to this exploration. Key cellular processes often depend on rare events, where a system must overcome energy barriers to spontaneously switch between metastable states. Extensively sampling those rare transitions, thus achieving ergodicity, is practically unfeasible for conventional, physically unbiased MD simulations. Even if we could dedicate enough computational power to this goal, we would produce massive, mostly redundant trajectory data of difficult interpretation—unless guided by prior knowledge. A more effective strategy is to carefully select the starting configurations of the simulations. In doing so, we split the problem of obtaining a long equilibrium trajectory into the more manageable task of sampling many shorter, meaningful paths. But how can we identify optimal starting configurations? Artificial intelligence (AI) offers an attractive solution. Using AI to both guide MD simulations and interpret the results is a core realization of the simulation intelligence paradigm and a foundational principle of this thesis. This work focused on both method development and applications to biologically relevant systems involving protein-protein and membrane-protein interactions. By combining rigorous path sampling with AI in an active learning framework, we demonstrated that accurate thermodynamic, kinetic, and mechanistic insights can be obtained with reasonable resource usage, even for processes occurring on time scales far beyond the reach of conventional MD simulations. In the first part of this work, we investigated the role of the amphipathic helix (AH) of the ATG3 protein in promoting LC3 lipidation on the phagophore membrane. Unsupervised machine learning (ML) revealed structural features that distinguish ATG3 AH from other membrane-sensing helices and guided the identification of an AH mutant that retains phagophore binding but compromises LC3 lipidation in vivo. Generative AI provided biologically relevant initial configurations for multiple unbiased MD simulations. Through supervised analysis and strategic reinitialization, we mapped the conformational landscape of the membrane-bound ATG3~LC3 complex and characterized the structural and dynamical differences between wild-type and mutant AHs. Based on these results, we proposed a speculative model describing the fine-tuned interplay between the AH, the lipids, and the rest of the complex. However, despite months of computation, the MD trajectories were not sufficiently long to extract comprehensive quantitative information, limiting our analysis to qualitative insights. For the remainder of this thesis, we addressed the sampling and interpretation challenges in systems shaped by rare-event transitions by adopting the AI for Molecular Mechanism Discovery (AIMMD) framework. At its core, AIMMD employs an iterative cycle where a ML model of the committor enhances path sampling simulations by orchestrating the selection of starting configurations—the shooting points (SPs)—from previously generated data. At the same time, the method uses the simulation results to train the model, further improving sampling efficiency. The committor function serves as the optimal reaction coordinate and, together with the sampled paths, can be analyzed to uncover the transition mechanism. What sets AIMMD apart from other AI-based enhanced sampling techniques is its minimal reliance on prior knowledge, its solid statistical mechanics foundations, its limited yet effective AI intervention in the simulations (still evolving from previous configurations under unbiased dynamics), and its strong emphasis on interpretability, all of which contribute to the method's robustness. In the second part of this work, we expanded the scope of AIMMD to obtain accurate free energy and transition rates estimates without the need for additional enhanced sampling simulations. To this end, we reintegrated all simulated excursions, previously considered failed attempts at generating transitions, in our sampling. We developed a data-efficient path reweighting algorithm incorporating free simulations around metastable states. Again, we relied on the committor model for optimal performance. We validated the updated method on the folding/unfolding of the mini protein chignolin, highlighting the computational advantages over standard MD simulations and the benefits of accessing the full equilibrium path ensemble. In the third part of this work, we demonstrated AIMMD's effectiveness in fully characterizing EGFR transmembrane dimerization and dissociation. In this way, we extended its applicability to a broader class of processes characterized by a single basin of attraction, where the transition is rare in only one direction. We emphasized the importance of dimer stability in triggering downstream signaling pathways and proposed AIMMD as a valuable tool for detecting stability changes under varying system conditions. Additionally, we further improved the method and its code implementation by parallelizing the computational effort, augmenting the committor training set, exploring path sampling alternatives to the original AI-guided TPS scheme, and streamlining the reweighting algorithm to improve its robustness and efficiency. In the fourth and final part of this work, we introduced an optimal rejection-free path sampling (RFPS) scheme by shifting our perspective to sampling a SP distribution. %an optimal path sampling scheme with no more rejected trials in the Markov chain. As a result, we could use all sampled paths for selecting SPs, enhancing exploration without compromising exploitation. RFPS integrates naturally into AIMMD, resulting in the updated RFPS-AIMMD algorithm, where free energy estimates are incorporated directly into the iterative cycle. RFPS-AIMMD improved committor learning through better training set coverage and demonstrated notable robustness on chignolin. This theoretical development bears broader implications for path sampling methods beyond AIMMD. Overall, the updated AIMMD framework offers a practical, interpretable, and computationally efficient alternative to long, unbiased MD simulations, with strong potential for further improvements and applications to more complex systems.

Gianmarco Lazzeri · 0 citations
Jul 2026

Learning Diffusion from Sparse Data: A Machine-Learning Bridge between Molecular Motion and Macroscopic Transport

Predicting molecular self-diffusion coefficients (D*) across chemical space remains challenging due to sparse experimental data and the high computational cost of molecular simulations. We present a data-centric machine learning framework that integrates experimental diffusion measurements with molecular dynamics simulations through a semisupervised distillation strategy. Unlike conventional approaches that treat limited experiments or simulations as direct ground truth, our method selectively incorporates simulation-derived D* only when they align with the uncertainty bounds of a Random Forest model initially trained on experimental D*. This enables controlled data set expansion, from 130 unique experimentally measured molecules to over 1,000 unique substances, while preserving label reliability. We further employ pretrained large language model embeddings to encode transferable chemical context beyond conventional descriptors, reducing predictive variance and improving generalization across chemically diverse systems. The resulting model achieves improved accuracy, with a held-out test set R2 of 0.87, an overall R2 of 0.92, and a test set mean absolute error of 0.13 × 10-9 m2 s-1 after iterative distillation across more than 1,000 chemically diverse molecules under near-ambient conditions (295-300 K). We further demonstrate that learned D* act as transferable, physically interpretable descriptors in skin permeability modeling, where their inclusion reduces test set error and improves correlation relative to models relying solely on conventional physicochemical descriptors. This work establishes a scalable framework for bridging sparse experimental measurements with broadly generalizable predictions, enabling interpretable and transferable modeling across distinct molecular transport phenomena.

Hojin Jung, Su-min Song, Sabari Kumar et al. · 0 citations

Related blog posts

GPT-Lab Aug 28, 2026

We built an AI factory for HVAC control

What does it take to trust AI-driven HVAC optimization? Our AI Model Factory combines agents, machine learning, reinforcement learning and deterministic checks in a governed workflow designed for messy, real-world building data. The post We built an AI factory for HVAC control appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.