This work releases OpenGEM26 (Open Generated Ensemble of Molecules, 2026), a large-scale dataset comprising 200,000 unique molecules and 4.4 million conformations composed of H, C, N, O, S and Cl with up to ten heavy atoms, providing a high-quality resource and robust ML potential for efficient simulations of sulfur- and chlorine-containing organic molecules.
Abstract
Density functional theory (DFT) serves as a reliable tool for atomistic molecular simulations, while machine learning potentials have become powerful complements to balance accuracy and efficiency. In this work, we release OpenGEM26 (Open Generated Ensemble of Molecules, 2026), a large-scale dataset comprising 200,000 unique molecules and 4.4 million conformations composed of H, C, N, O, S and Cl with up to ten heavy atoms. All calculations are carried out at the {\omega}B97X-D/Def2-SVP and Def2-TZVP levels with dispersion corrections, and complete structural optimization trajectories and abundant non-equilibrium structures are recorded. Statistical analyses confirm that this dataset covers a broader conformational space than QM9 in terms of energy, bond lengths and bond angles. A graph neural network-based potential GPTFF-mol is trained using the new dataset, achieving an energy mean absolute error of 16 meV/molecule, which is equivalent to 0.82meV/atom, and superior force prediction performance compared with ANI-2x. Validated by butane rotation and keto-enol tautomerization tests, the model accurately describes molecular dynamical behaviors and reaction barriers at distorted geometries. This work provides a high-quality resource and robust ML potential for efficient simulations of sulfur- and chlorine-containing organic molecules.
Molecular net force provides a complementary diagnostic of force consistency but should not be interpreted as a direct measure of atomic-force accuracy, and the proposed framework is readily applicable to quality assessment of DFT datasets used for machine learning.
H. Assem, K. Appiah, C. Subaar· Journal of Molecular Graphic...· 0 citations
The SBMR-CNN model demonstrates highly competitive accuracy, outperforming the CM, Uni-Mol+, and MPNN-2D benchmarks, while closely approaching the performance of the more computationally intensive MPNN-3D and SOAP descriptors, as well as the RF-MF model.
Abdulaziz W. Alherz, C. Tezak, Mohammed S. Alhajeri· Industrial & Engineering...· 0 citations
This work lays the foundation for NNPs where solvation is an integral part of the model, enabling the development of multiscale NNPs for simulating large biomolecular systems.
Moritz Thürlemann, Felix Pultar, Igor Gordiy et al.· Scientific Data· 0 citations
DensIP is introduced, a physics-based model of intermolecular interactions that uses machine-learned electron densities and only four universal parameters that outperforms state-of-the-art general-purpose MLFFs for long-range interactions and can be applied to molecules as large as drug ligands.
Dahvyd Wing, Mihail Bogojeski, Szabolcs Góger et al.· 0 citations
Coupled-cluster theory defines the accuracy standard for molecular electronic-structure properties but scales too steeply for routine application, whereas density-functional theory is affordable yet systematically biased. We resolve this trade-off with a single equivariant network, MEHnet-MG, that predicts an effective one-electron Hamiltonian from one inexpensive B3LYP/def2-SVP calculation and derives a broad suite of properties from it (energy, optical gap, dipole, quadrupole, polarizability, Mulliken atomic charges, and Mayer bond orders) at coupled-cluster accuracy across nine main-group elements, including the under-served phosphorus, sulfur, and chlorine chemistries. The model is trained on a new in-house dataset of multi-property labels computed at the CCSD(T) level for all nine elements. On a held-out test set, it reduces the error of every property by a factor of 3.8 to 230 relative to semi-local, hybrid, and double-hybrid DFT (referenced to composite CCSD(T)/cc-pVTZ; Methods), while adding only ~25 ms wall time per molecule, delivering coupled-cluster-quality predictions at the cost of a single DFT calculation. Critically, deriving every property from a predicted Hamiltonian rather than pooling per-atom features builds the correct size-scaling into the model architecture: on pi-conjugated oligothiophenes it matches finite-field CCSD polarizability and the EOM-CCSD optical gap to ~2% at the largest sizes where those references remain affordable (44 and 37 atoms, where a single CCSD field point already costs ~500x the model's entire inference) and extrapolates the corrected trends to 58-atom chains, a regime where pooling-based architectures fail by construction. Accurate extrapolation is therefore set by the model's inductive bias rather than by the training data.
A structure-aware graph neural network is trained to predict cross-functional energy residuals and align inconsistent DFT energy scales, which enables reliable predictions of phase stability, battery voltage profiles, and reaction thermodynamics, while allowing the integration of multi-source DFT data to advance the development of high-performance materials foundation models.
Yidong Huang, Tenglong Lu, Hanwen Kang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.