Skip to content

Similar papers

Preprint Jul 2026

Anisotropic representations for E(3)-equivariant machine learning coarse-grained potentials

Coarse-graining (CG) lowers the computational cost of atomistic simulations by representing groups of atoms as effective interaction sites, reducing the degrees of freedom of the system but often compromising structural fidelity or requiring system-specific parameterization. Here, we introduce a novel anisotropic machine learning CG potential that extends the point particle representation of atomic nuclei to massive ellipsoidal beads with orientation-dependent features, enabling the learning of energies, forces, and torques directly from atomistic data. The anisotropic representation is physically motivated for polar and asymmetric molecules, where directional interactions and shape anisotropy play important roles in determining structure and dynamics. Using an equivariant message-passing neural network, the model accurately reproduces radial and angular distribution functions as well as relative orientation correlations in liquid water, demonstrating that both translational and rotational dynamics are well captured. Comparison with an isotropic baseline reveals that the lack of orientation information leads to systematic errors in short and long range order and degradation of angular correlations, proving orientation features are essential for accurate coarse-graining. The anisotropic model also exposes rotational structural observables fundamentally inaccessible to isotropic representations, with minimal computational overhead. Even for coarse-graining just three degrees of freedom, CG simulations achieve 7-27$\times$ speedups while preserving structural fidelity, highlighting the efficiency gains of this systemic reduction. This framework establishes the feasibility and necessity of learned equivariant representations for anisotropic CG modeling and provides a path towards accurate and efficient mesoscopic simulations of complex molecular liquids, polymers, and biomolecular systems.

V. Shankar, Emil Annevelink · 0 citations
Aug 2026

Molecular Property Prediction via Sparse Binary Matrix Representation and Convolutional Neural Networks

A simple and interpretable matrix-based representation is presented for predicting molecular properties, specifically individual HOMO and LUMO frontier orbital energies and their resulting energy gaps, of functionalized organic molecules using a Convolutional Neural Network (CNN). Each molecule is encoded as a sparse binary matrix (SBMR) that captures the identity and position of substituents on a fixed molecular backbone. The model was initially benchmarked across four molecular families: n-butane, i-butane, cyclobutadiene, and quinone, achieving a combined RMSE of 4.0 kcal mol–1 for gap predictions compared to DFT-computed references, with over 85% of predictions falling within ±5% error. To contextualize this performance, the model was benchmarked against six established featurization methods spanning 2D topology and 3D physics-based approaches: the Coulomb Matrix (CM), Smooth Overlap of Atomic Positions (SOAP), 2D and 3D Message-Passing Neural Networks (MPNN), Random Forest with Morgan Fingerprints (RF-MF), and Uni-Mol+. The SBMR-CNN model demonstrates highly competitive accuracy, outperforming the CM, Uni-Mol+, and MPNN-2D benchmarks, while closely approaching the performance of the more computationally intensive MPNN-3D and SOAP descriptors, as well as the RF-MF model. This is achieved while offering distinct advantages through a dramatically smaller feature space and less stringent input data requirements. To demonstrate extensibility to complex catalytic systems, the architecture was applied to a combinatorial data set of 1,4-dihydropyridine derivatives, a class of redox mediators utilized in electrochemical and biochemical applications. For these highly functionalized heterocycles, the model successfully decoupled the energy gap into its constituent levels, predicting HOMO and LUMO energies with an RMSE of 2.8 and 2.5 kcal mol–1, respectively. The resulting framework couples high predictive accuracy with representational interpretability, offering a transparent and customizable tool for property prediction with direct applications in molecular screening, rational design, and electrocatalyst optimization.

Abdulaziz W. Alherz, C. Tezak, Mohammed S. Alhajeri · 0 citations
Open access Jan 2026

An algebraic graph neural network model for protein-ligand binding affinity prediction

An Algebraic Graph Neural Network model designed to encode molecular structures into a low-dimensional graph representation while preserving critical biochemical interactions is introduced, demonstrating superior performance in binding affinity prediction compared to state-of-the-art scoring functions.

Augustine Ouru, Xi Chen, Cameron Yeagle et al. · 0 citations
Preprint Jul 2026

Rem3Di: Learning smooth, chiral 3D molecular descriptors from atomistic foundation models

Foundation machine-learned interatomic potentials (MLIPs) are trained on large quantum-mechanical datasets and generalise across broad regions of chemical and configurational space. Beyond their usual role in accelerating sampling-based simulations, their internal representations encode chemically rich local atomic environments. Here, we introduce Rem3Di, a representation-learning framework that repurposes latent features from atomistic foundation models as transferable molecular descriptors for property prediction and virtual screening. Rem3Di combines a potential's per-atom features into a single fixed-length descriptor of the whole molecule that varies smoothly with three-dimensional structure and is invariant to the ordering of the atoms. The descriptor can be used directly or fine-tuned for specific prediction tasks. To capture molecular handedness, Rem3Di constructs pseudoscalar features, which are unchanged by rotation but reverse sign under mirror reflection. This lets the descriptor distinguish enantiomers, which can differ in activity and toxicity. The transformer is pretrained on large molecular datasets by reconstructing corrupted atom features, so no experimental labels are required. Across public drug-property benchmarks, Rem3Di matches or exceeds published baselines without relying on classical 2D fingerprints. Additionally, the same descriptor yields chemically meaningful differentiation of transition-metal complexes without predefined bonding rules or handcrafted representations. Rem3Di therefore provides a route from simulation-trained atomistic representations to transferable, chirality-aware molecular representations for chemical machine learning.

Steffen Wedig, Felix Burton, Rokas Elijošius et al. · 0 citations
Conference Jul 2026

GDGraph: Geometry-Enhanced Dual-View Graph for Molecular Representation Learning

Learning effective molecular representations is crucial for accurate property prediction in AI-aided drug discovery. However, most existing molecular pre-training methods are still primarily based on 2D topological graphs, limiting their ability to exploit 3D geometric information. Moreover, methods that do incorporate 3D geometry often do not distinguish between the roles of atom-centered and bond-centered representations. To address these limitations, we propose GDGraph, a geometryenhanced dual-view framework for molecular representation learning. GDGraph models molecular geometry from two complementary structural perspectives: an atom view for capturing global spatial dependencies and a bond view for modeling local geometric patterns. To support this dual-view design, we introduce a multi-scale geometric feature encoding scheme and a view-specific geometry-aware learning strategy, enabling each view to focus on the geometric dependencies it is best suited to capture. Extensive experiments demonstrate that GDGraph achieves strong and stable performance on molecular property prediction benchmarks, and effectively predicts geometrysensitive quantum chemical properties on the QM9 dataset.

Yu Liu, Jonathan D. Hirst, Jianfeng Ren et al. · 0 citations