Back to feed
Preprint

From oligomers to entangled polymers: How to train a transferable machine learning interatomic potential

Aug 2026 · 0 citations · 67 references
Physics

TL;DR

This work investigates several aspects of developing MLIPs for polymers, utilizing polyethylene as a representative, yet simple model system, and finds that the ACE potential accurately reproduces key thermodynamic, structural and dynamical properties.

Abstract

Over the past decade, Machine Learning Interatomic Potentials (MLIPs) have emerged as a powerful technique for performing molecular dynamics (MD) simulations with nearly ab initio accuracy. Alongside the development of new descriptors and advanced machine learning architectures, sophisticated procedures for the generation of diverse and accurate reference datasets have been established. To date, research has focused primarily on MLIPs for crystalline or amorphous inorganic and small molecular systems; however, large macromolecules such as polymers remain underrepresented in the literature, despite beeing an important class of materials. In this work, we investigate several aspects of developing MLIPs for polymers, utilizing polyethylene as a representative, yet simple model system. First, we compare various local atomic descriptors, identifying the Atomic Cluster Expansion (ACE) as the most effective for this application. Second, we implement and automatized active learning scheme to efficiently generate diverse training data and demonstrate that ACE potentials fitted on small oligomers are transferable to larger polymers. Given that the accurate reproduction of the density depends critically on a correct description of intermolecular interactions, which are far more complex to learn than intramolecular interactions, we carefully evaluate the performance of the ACE potentials with respect to non-bonded interactions. By utilizing the computationally efficient OPLS-AA force field as a ground truth reference, we are able to perform a direct comparison of nanosecond-scale MD trajectories resulting from the ACE and reference potential. We find that the ACE potential accurately reproduces key thermodynamic, structural and dynamical properties.

View source

Similar papers

Preprint Aug 2026

Data-Efficient Construction of Material-Specific Machine-Learning Interatomic Potentials from Ab Initio Molecular Dynamics Trajectories

Pretrained machine-learning interatomic potentials, so-called universal or foundation models offer an appealing starting point for atomistic simulations, but their accuracy for material-specific observables often remains limited without additional reference data (fine-tuning). Here, we systematically quantify how much first-principles data are required to convert universal models into ab initio-accurate material-specific potentials, and ask whether fine-tuning is necessarily preferable to training from scratch. We compare five universal MLIP frameworks, MACE-MP-0, SevenNet-0, GRACE-1L-OAM, MatterSim-v1-5M and ORB-v2, across seven chemically diverse systems incorporating rare and reactive events. Fine-tuning on only 10 AIMD-derived configurations is insufficient for the investigated systems; 200 configurations succeed in favorable cases, but the outcome remains strongly system-dependent. By contrast, 2000 AIMD configurations constitute a robust default, yielding low force and energy errors and reproducing the target material-specific observables. Moderately dense sub-sampling of the AIMD trajectory reduces the required trajectory length tenfold with little loss in model quality. Training from scratch on the same datasets is competitive with, and often slightly more accurate than, naive fine-tuning for MACE and SevenNet, whereas GRACE requires more data. The energy profile for a sulfur-vacancy jump in MoS$_2$ reveals that low trajectory-level errors do not guarantee a correct reaction profile, highlighting the need for observable-level validation. Finally, we show that averaging independently trained models improves predictions in scarce-data regimes at no additional first-principles cost. Together, these results provide practical guidelines for converting limited AIMD reference data into reliable material-specific MLIPs for nanosecond-timescale simulations at near-DFT accuracy.

Jonas Hänseroth, Christian Dressler · 0 citations
Preprint Jul 2026

Fast and Accurate Foundation Models for Equivariant Machine-Learned Interatomic Potentials

Machine-learned interatomic potentials (MLIPs) have emerged as a transformative tool for computational materials science and chemistry, with universal potentials trained on large and diverse datasets now routinely deployed as'foundation models'for downstream fine-tuning in targeted chemical spaces. Many scientific applications of the resulting models, such as molecular dynamics (MD), require high inference and training speeds as well as accuracy. In this work, we examine the limits of equivariant MLIPs, which directly encode physical symmetries in model architectures, to achieve these competing targets -- particularly in the regime of extremely large datasets where data efficiency is less critical. We show how this trade-off can be addressed, and present a family of foundation potentials in the NequIP and Allegro equivariant MLIP architectures which achieve leading inference speeds and strong scalability as well as excellent accuracies across a range of community benchmarks -- spanning materials discovery, thermal conductivity prediction, and near-equilibrium mechanical and thermodynamic properties. Accelerations implemented within the NequIP infrastructure now permit training of high-accuracy foundation potentials on ultra-large datasets with dramatically reduced computational cost. Alongside, we show that efforts to improve model accuracy for materials discovery should focus on dataset diversity and improved, consistent descriptions of transition metal compound energy surfaces.

Seán R. Kavanagh, Chuin Wei Tan, Menghang Wang et al. · 0 citations
Jun 2026

ElemeNet: Multiscale Molecular Machine Learning with Uncertainty Quantification Across the Periodic Table

The ElemeNet software package enables the training of advanced ML models for diverse properties and datasets with an enlarged range of elemental compositions, and introduces moiety predictions, a unified, general-purpose software package for molecular machine learning.

Jacob W. Toney, S. Darouich, Yiran Wang et al. · 0 citations
Preprint Jul 2026

Amorphous materials as a frontier challenge for universal interatomic potentials

Pre-trained or'foundational'machine-learned interatomic potentials (MLIPs) are now widely used in materials modelling. However, early pre-trained models and benchmarks have largely focused on ordered, crystalline structures, and their transferability to non-crystalline solids remains unclear. Here, we show that the amorphous state is indeed a central challenge for future universal MLIPs, based on a systematic evaluation of current mainstream models in this domain. We introduce a benchmarking framework built on a curated reference dataset of canonical amorphous systems, as well as validation for structures and properties. Our study identifies limitations in the transferability of many current pre-trained models and investigates fine-tuning strategies tailored to disordered phases. Together, our results can facilitate future applications of MLIPs in the fast-growing field of amorphous functional materials, and they provide guidance for designing next-generation training datasets and transferable atomistic models.

Natascia L. Fragapane, Volker L. Deringer · 2 citations
Aug 2026

ALF: Open-Source Active Learning Framework for Atomistic Modeling

Machine learning interatomic potentials (MLIPs) have surged in popularity over the last two decades, with many model architectures now openly available. As data-driven models, MLIPs critically depend on high-fidelity (i.e., physically accurate) training data produced by electronic structure calculations. However, assembling large and chemically diverse datasets can be a complex and time-consuming endeavor, often requiring the manual selection of representative atomic configurations and the execution of hundreds to millions of electronic structure simulations. To address this challenge, we introduce the Active Learning Framework (ALF), an open-source Python package designed to streamline the design and deployment of MLIP training datasets on High Performance Computing resources. ALF automatically selects new configurations from undersampled regions of the potential energy surface where the MLIP exhibits high uncertainty, schedules electronic structure calculations across available computational resources, and retrains MLIPs on the fly, thereby reducing manual intervention and limiting human bias. As a demonstration, we applied ALF to generate an actively learned dataset for molten salt mixtures consisting of F, Li, Na, Be, and K atoms. An MLIP trained on this data was then employed to predict melting point, viscosity, density, radial distribution function, and specific heat, which are computationally resource-intensive to evaluate via first-principles molecular dynamics. These results were subsequently validated against experimental data. Collectively, these findings illustrate ALF’s effectiveness in compiling datasets that capture essential chemical and structural regimes, thereby virtually eliminating manual curation.

V. Grizzi, P. Lohr, Nikita Fedik et al. · 0 citations
Open access Feb 2026

Machine learning of electronic structure and atomistic properties from the external potential.

Electronic structure calculations remain a major bottleneck in atomistic simulations and, not surprisingly, have attracted significant attention in machine learning (ML). Most existing approaches learn a direct map from molecular geometries, typically represented as graphs or encoded local environments, to molecular properties or use ML as a surrogate for electronic structure theory by targeting quantities, such as Fock or density matrices expressed in an atomic orbital (AO) basis. Inspired by the Hohenberg-Kohn theorem, in this work, we propose an operator-centric framework in which the external (nuclear) potential, expressed in an AO basis, serves as the model input. From this operator, we construct hierarchical, body-ordered representations of atomic configurations that closely mirror the principles underlying several popular atom-centered descriptors. At the same time, the matrix-valued nature of the external potential provides a natural connection to equivariant message-passing neural networks. In particular, we show that successive products of the external potential provide a scalable route to equivariant message passing and enable an efficient description of nonlocal effects. We demonstrate that this approach can be used to model molecular properties, such as energies and dipole moments, from the external potential or to learn effective operator-to-operator maps, including mappings to the Fock matrix from which multiple molecular observables can be simultaneously derived.

Jigyasa Nigam, T. Smidt, G. Dusson · 2 citations