Back to #artificial intelligence

Coupled-cluster molecular properties across the main group that extrapolate beyond training size

Aug 2026 · 0 citations · 33 references
Physics Computer Science

Abstract

Coupled-cluster theory defines the accuracy standard for molecular electronic-structure properties but scales too steeply for routine application, whereas density-functional theory is affordable yet systematically biased. We resolve this trade-off with a single equivariant network, MEHnet-MG, that predicts an effective one-electron Hamiltonian from one inexpensive B3LYP/def2-SVP calculation and derives a broad suite of properties from it (energy, optical gap, dipole, quadrupole, polarizability, Mulliken atomic charges, and Mayer bond orders) at coupled-cluster accuracy across nine main-group elements, including the under-served phosphorus, sulfur, and chlorine chemistries. The model is trained on a new in-house dataset of multi-property labels computed at the CCSD(T) level for all nine elements. On a held-out test set, it reduces the error of every property by a factor of 3.8 to 230 relative to semi-local, hybrid, and double-hybrid DFT (referenced to composite CCSD(T)/cc-pVTZ; Methods), while adding only ~25 ms wall time per molecule, delivering coupled-cluster-quality predictions at the cost of a single DFT calculation. Critically, deriving every property from a predicted Hamiltonian rather than pooling per-atom features builds the correct size-scaling into the model architecture: on pi-conjugated oligothiophenes it matches finite-field CCSD polarizability and the EOM-CCSD optical gap to ~2% at the largest sizes where those references remain affordable (44 and 37 atoms, where a single CCSD field point already costs ~500x the model's entire inference) and extrapolates the corrected trends to 58-atom chains, a regime where pooling-based architectures fail by construction. Accurate extrapolation is therefore set by the model's inductive bias rather than by the training data.

View source

Similar papers

Preprint Jul 2026

Learning to Converge: Warm-Starting DFTB Self-Consistent Charges with Machine Learning

Semiempirical electronic structure methods such as Density-Functional Tight-Binding (DFTB) offer a computationally efficient approach to molecular and materials simulations, bridging the gap between first-principles accuracy and classical force field speed while retaining full access to electronic properties. However, DFTB calculations based on self-consistent charge (SCC) schemes can still suffer from slow convergence, particularly for complex molecular and materials systems, making the iterative procedure a significant bottleneck in large-scale simulations and high-throughput workflows. We present a machine learning approach that accelerates DFTB simulations by predicting optimal initial atomic charges. Using element-specific models based on the Smooth Overlap of Atomic Positions descriptor and kernel ridge regression, we train charge models on reference calculations and demonstrate that ML-predicted initial charges consistently and significantly improve SCC convergence across diverse chemical systems including organic molecules, biomolecules, water clusters, transition metal oxides and solid electrolytes.

Maximilian L. Ach, Karsten Reuter, C. Panosetti · 0 citations
Open access Jul 2026

Reactive Chemistry at the Unrestricted Coupled Cluster Level: High-Throughput Calculations for Training Machine Learning Potentials

Modeling chemical reactions accurately at the atomistic level requires high-level electronic structure theory due to the presence of unpaired electrons and the need to properly describe the energetics of bond breaking and bond formation. Commonly used approaches such as density functional theory (DFT) frequently fail for this task due to deficiencies that are well recognized. However, for high-fidelity approaches, creating large data sets of energies and forces for reactive processes to train machine learning interatomic potentials (MLIPs) or force fields is daunting. For example, the use of the unrestricted coupled cluster level of theory has previously been seen as unfeasible due to high computational costs, the lack of analytical gradients in many computational codes, and additional challenges such as constructing suitable basis set corrections for forces. In this work, we develop new methods and workflows to overcome the challenges inherent to automating unrestricted coupled cluster calculations. Using these advancements, we create a data set of gas-phase reactions containing energies and forces for 3119 different organic molecules configurations calculated at the gold-standard level of unrestricted CCSD(T) (coupled cluster singles doubles and perturbative triples). With this data set, we provide an analysis of the differences between the density functional and unrestricted CCSD(T) descriptions. We develop a transferable MLIP for gas-phase reactions, trained on unrestricted CCSD(T) data, and demonstrate the advantages of transitioning away from DFT data. Transitioning from training to DFT to training to UCCSD(T) data sets yields an improvement of more than 0.1 eV/Å in force accuracy and over 0.1 eV in activation energy reproduction.

Alice E. A. Allen, Rui Li, Sakib Matin et al. · 0 citations
Open access Jul 2026

Sparse Linear Surrogates Match Neural Network Potentials on the SPICE Biomolecular Benchmark with Three Orders of Magnitude Smaller Training Sets

We introduce the orbital cluster expansion (OCE), a linear regression on physics-motivated local features derived from atomic orbital eigenenergies, and benchmark it against the SPICE 2.0 biomolecular data set at the ωB97M-D3BJ/def2-TZVPPD level. With regression of formation energies on 677 dipeptides spanning the natural amino acids, ridge regression on 414 OCE features attains a parent-stratified test root-mean-square error of 30 meV per atom with Spearman ρ = 0.97 and R 2 = 0.95 against a target spread of only 0.13 eV per atom, matching MACE-OFF23(L) and ANI-2x trained with 104–106 conformations but with ∼103 fewer training points. Comparable accuracy holds on 500 PubChem drug-like molecules and 500 DES370K dimers. We characterize a fundamental dual regime: intermolecular ranking is preserved across chemistries, while intraconformer ranking is random because the basis cannot resolve geometry-only variation within a fixed connectivity. OCE is a transparent, physically interpretable surrogate for intermolecular biomolecular screening.

D. L. Azevedo · 0 citations
Open access Feb 2026

Machine learning of electronic structure and atomistic properties from the external potential.

Electronic structure calculations remain a major bottleneck in atomistic simulations and, not surprisingly, have attracted significant attention in machine learning (ML). Most existing approaches learn a direct map from molecular geometries, typically represented as graphs or encoded local environments, to molecular properties or use ML as a surrogate for electronic structure theory by targeting quantities, such as Fock or density matrices expressed in an atomic orbital (AO) basis. Inspired by the Hohenberg-Kohn theorem, in this work, we propose an operator-centric framework in which the external (nuclear) potential, expressed in an AO basis, serves as the model input. From this operator, we construct hierarchical, body-ordered representations of atomic configurations that closely mirror the principles underlying several popular atom-centered descriptors. At the same time, the matrix-valued nature of the external potential provides a natural connection to equivariant message-passing neural networks. In particular, we show that successive products of the external potential provide a scalable route to equivariant message passing and enable an efficient description of nonlocal effects. We demonstrate that this approach can be used to model molecular properties, such as energies and dipole moments, from the external potential or to learn effective operator-to-operator maps, including mappings to the Fock matrix from which multiple molecular observables can be simultaneously derived.

Jigyasa Nigam, T. Smidt, G. Dusson · 2 citations
Book Open access Aug 2026

UniHam: A Large-Scale SOC-Complete Dataset and Benchmark for Hamiltonian Learning in Materials

Accurate prediction of electronic Hamiltonians would enable broad property inference while avoiding the high computational cost of Density Functional Theory (DFT). However, progress toward general-purpose materials foundation models is limited by a data bottleneck: existing Hamiltonian datasets are typically small, lack structural diversity, and often omit essential relativistic physics such as spin--orbit coupling (SOC). We therefore construct UniHam, a large-scale Hamiltonian dataset and benchmark suite comprising 100,000+ DFT-computed complex-valued Hermitian Hamiltonians with full SOC, covering 72 elements and a wide range of crystal geometries and symmetries (spanning diverse lattice types and space-group families). Building on UniHam, we benchmark two representative state-of-the-art models under a standardized protocol and introduce complementary evaluation metrics that jointly assess three dimensions: (i) Hamiltonian reconstruction accuracy, (ii) out-of-distribution (OOD) generalization across composition/symmetry shifts, and (iii) the ability to support downstream property prediction from the predicted Hamiltonians. Experiments on UniHam demonstrate that the proposed benchmark and metrics effectively differentiate model capabilities, revealing intrinsic SOC- and element-dependent failure modes, large variations in compositional OOD robustness, and the necessity of spectral-level evaluation to assess whether Hamiltonian predictions reliably support downstream electronic-structure properties. Overall, UniHam provides a reproducible, SOC-complete benchmark that can sharpen model comparisons and accelerate the development of next-generation foundation models for quantum materials.

Yuewen Huang, Pin Chen, Yutong Lu · 0 citations
Preprint Jul 2026

Implementations of Quantum and Classical Topology-Aligned Architectures for Molecular Property Prediction

For low-data and resource-constrained regimes typical of quantum chemistry, parameter-efficient learning is a key objective. Here, we propose a topology-aligned inductive bias in which the model architecture mirrors the molecular bond graph: atoms map to a fixed register of computational units, and bonds determine which pairs interact through shared learnable parameters. This principle is instantiated in two architectures: a variational quantum circuit (Iso-QGNN) and a parameter-matched classical message-passing network (Iso-CGNN). The models are benchmarked on HOMO-LUMO and dipole moment binary classification tasks over the QM9 benchmark. With 64 trainable parameters, the implementations achieve test AUCs of approximately 0.89 (quantum) and 0.92 (classical) on the gap task, and close to 0.78 (both) on the dipole task. The models reach 90% of asymptotic performance within about 300 training molecules and gradient norms remain stable throughout training. These results indicate that the topology-aligned inductive bias is the active ingredient driving parameter efficiency at QM9 scale, with implications for matched-baseline benchmarking in quantum machine learning.

James T. Pegg, H. Valencia, Ronin Wu · 0 citations

Related blog posts