Oct 2025· Journal of Chemical Theory and Computation· Vol 22, pp. 1613 - 1620· 7 citations· 43 references
MedicinePhysics
TL;DR
A committor-based method that promotes frequent transitions between the metastable states of the system and allows extensive sampling of the process transition state ensemble and highlights the advantages of a graph-based approach in describing the role of solvent molecules in systems, such as ion pair dissociation or ligand binding.
Abstract
The study of rare events is one of the major challenges in atomistic simulations, and several enhanced sampling methods toward its solution have been proposed. Recently, it has been suggested that the use of the committor, which provides a precise formal description of rare events, could be of use in this context. We have recently followed up on this suggestion and proposed a committor-based method that promotes frequent transitions between the metastable states of the system and allows extensive sampling of the process transition state ensemble. One of the strengths of our approach is being self-consistent and semiautomatic, exploiting a variational criterion to iteratively optimize a neural-network-based parametrization of the committor, which uses a set of physical descriptors as input. Here, we further automate this procedure by combining our previous method with the expressive power of graph neural networks, which can directly process atomic coordinates rather than descriptors. Besides applications on benchmark systems, we highlight the advantages of a graph-based approach in describing the role of solvent molecules in systems, such as ion pair dissociation or ligand binding.
Many molecular dynamics simulations aim at studying transitions between two states (from reactants to products). In this context, the committor function (which gives for a given molecular configuration the probability to reach the product state before the reactant state) is a pivotal quantity, in particular because it is the optimal importance function for rare event simulation methods such as importance sampling or splitting techniques. These methods are used to sample the reactive path ensemble, and estimate for example the transition rate. However, learning such a function is generally a challenging task due to the high dimensionality of the configuration space. In this work, after reviewing the existing methodologies to construct approximate committor functions, a new loss function based on the application of It\={o}'s formula is proposed to learn the committor function with a minimization procedure on the parameters of a neural network. After comparing this novel approach to existing procedures on the M\"uller--Brown potential, we introduce a coupling strategy with the Adaptive Multilevel Splitting method to better approximate the committor function using a better sampling of the reactive trajectories. This methodology in which the committor function is iteratively learned only requires initially the knowledge of the reactant and product states.
Thomas Pigeon, Gabriel Stoltz, T. Lelièvre· 0 citations
Mapping an atomic structure to a compact set of geometric descriptors is an essential step in any machine-learning application to atomic-scale modeling. A powerful and widely-used approach can be understood as a discretization of the histogram of pair distances, triangles, etc., that results in a hierarchy of symmetry-invariant atom-centered descriptors. Unfortunately, the lower rungs on this hierarchy (two, three, four-neighbor clusters) were found to be incomplete, with symmetry-unrelated pairs of structures having exactly the same descriptors. However, all the ``descriptor degeneracies''reported so far are resolved by considering larger clusters of neighbors to build the descriptors. We report examples of 3D structures that are indistinguishable even if one considers clusters of up to seven neighbors, and to arbitrary order when considering a practical level of discretization of the descriptors, discovered with the assistance of large language models. The key ingredients in their construction can be traced to results that have been known for decades in different communities; the model was able to find the references and recognize their significance for the problem at hand. We believe this experiment exposes an extremely fruitful usage pattern for AI in science: translating results between different communities and application domains, accelerating the process by which serendipitous discoveries in a field become paradigm-shifting breakthroughs in another.
M. Domina, Michele Ceriotti· arXiv.org· 1 citation
This work proposes an operator-centric framework in which the external (nuclear) potential, expressed in an AO basis, serves as the model input and builds hierarchical, body-ordered representations of atomic configurations that closely mirror the principles underlying several popular atom-centered descriptors.
Jigyasa Nigam, T. Smidt, G. Dusson· Journal of Chemical Physics· 2 citations
Abstract.
Collective variables (CVs) play a crucial role in capturing rare events in high-dimensional systems, motivating the continual search for principled approaches to their design. In this work, we revisit the framework of quantitative coarse graining and identify the orthogonality condition from Legoll and Lelievre (2010) as a key criterion for constructing CVs that accurately preserve the statistical properties of the original process. We establish that satisfaction of the orthogonality condition enables error estimates for both relative entropy and pathwise distance to scale proportionally with the degree of scale separation. Building on this foundation, we introduce a general numerical method for designing neural network–based CVs that integrates tools from manifold learning with group-invariant featurization. To demonstrate the efficacy of our approach, we construct CVs for butane and achieve a CV that reproduces the antigauche transition rate with less than ten percent relative error and within two standard deviations of the empirically measured rate. Additionally, we provide empirical evidence challenging the necessity of uniform positive definiteness in diffusion tensors for transition rate reproduction and highlight the critical role of light atoms in CV design for molecular dynamics.
Shashank Sule, Arnav Mehta, Maria K. Cameron· Multiscale Modeling & Si...· 0 citations
Adaptive sampling accelerates the exploration of conformational space in molecular dynamics (MD) simulations by repeatedly analyzing the accumulated trajectories and seeding a new round of simulations from informative configurations. A growing collection of adaptive sampling policies has been proposed, each built around a particular notion of what makes a configuration informative, yet these methods are scattered across separate and often incompatible implementations, which complicates their systematic comparison and their combined use in meta adaptive sampling schemes. Here, we present AdaptivePy, a compact and extensible Python framework that implements nine seed-selection policies behind a single configuration-driven interface, spanning simple population-based baselines, several established machine-learning and geometry-based methods, and two ensemble or meta sampling policies introduced in this work. We show that the shared implementation reproduces the characteristic selection behavior of each policy on a series of analytic benchmark landscapes. We also introduce a new adaptive sampling scheme that employs TS-DAR, a deep learning framework originally designed to identify transition states, into an acquisition criterion that drives the discovery of an entire multi-basin landscape starting from a single basin. We further demonstrate that the common interface enables meta adaptive sampling policies, which aggregate the rankings of several policies into a single set of seeds. AdaptivePy thereby provides a unified testbed for the adoption, benchmarking, and continued development of adaptive sampling methods for biomolecular MD simulations.
Hassan Nadeem, D. Kleiman, Diwakar Shukla· bioRxiv· 0 citations
A notably simple procedure, a method the authors refer to as deletions, yields superior performance over an array of alternative extraction methods for extracting atomic environments from large, bulk configurations and embedding them into smaller configurations suitable for DFT calculations with periodic boundary conditions.
Jared Stimac, Fei Zhou, Kyle Bushick et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 2, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.