Skip to content
Preprint

Evaluating Electrostatic Embedding MLIP/MM for Relative Binding Free Energy Calculations

Aug 2026 · 0 citations · 83 references
Physics

TL;DR

Electrostatic embedding improved every accuracy and correlation metric for TYK2 but performed comparably to the classical and mechanical-embedding baselines for CDK2, thrombin, p38 and JNK1, and standard single-molecule energy and charge benchmarks were not good predictors of this target-dependent outcome.

Abstract

Alchemical relative binding free energy (RBFE) calculations are limited by the fixed-charge approximation of classical force fields. Hybrid machine learning interatomic potential/molecular mechanics (MLIP/MM) schemes correct ligand strain, but under mechanical embedding still describe ligand--environment electrostatics with static point charges. Electrostatic embedding schemes coupling machine-learned charges to the MM environment have been proposed and validated against QM/MM for simple systems, but not tested in a production alchemical workflow. We take the electrostatic embedding scheme of Semelak et al.\ and evaluate it on protein--ligand RBFE. We trained a TensorNet2 model, \texttt{AceFF-2-RESP-1}, on $10^{6}$ conformations from the AceFF dataset, jointly predicting energies, forces and Restrained Electrostatic Potential (RESP) charges. We chose RESP over MBIS for commensurability with the AMBER-family force field it couples to. The predicted charges enter the short-range direct-space part of the particle mesh Ewald sum, with Thole damping to prevent polarization catastrophes during alchemical transformations. We tested the scheme across five targets from the Wang et al.\ benchmark set, fixed in advance by a prior study, with three replicates per edge and matched protocols. Electrostatic embedding improved every accuracy and correlation metric for TYK2 ($\Delta\Delta G$ RMSE $0.86 \rightarrow 0.45$~kcal/mol against GAFF2), but performed comparably to the classical and mechanical-embedding baselines for CDK2, thrombin, p38 and JNK1. Standard single-molecule energy and charge benchmarks were not good predictors of this target-dependent outcome. TYK2 combined good $\Delta\Delta G$ accuracy with the lowest force error on the Schr\"odinger benchmark, but this pattern did not hold for the other targets.

View source

Similar papers

Jul 2026

Charged Systems in Absolute Binding Free Energy Calculations: An Analytical Electrostatic Approach.

Alchemical free energy perturbation (FEP) is one of the most rigorous methods for predicting protein-ligand binding affinities, yet charged-ligand calculations suffer from finite-size electrostatic artifacts introduced by periodic boundary conditions, which can bias results by several kcal·mol-1. Existing approaches each have limitations: finite-size correction methods rely on approximate dielectric models and Poisson-Boltzmann (PB) calculations, while alchemical co-ion methods introduce alchemically transformed particles, causing spurious interactions and sampling difficulties. Here we present Electrostatic Interaction Decoupling (EID), a postprocessing approach that combines an exact algebraic isolation of the ligand-environment linear electrostatic interaction under the neutral-environment condition with an analytical correction for the residual periodic-boundary offset. By separating the physical ligand-environment interaction from artifact-contaminated terms, EID corrects charge-changing FEP results without PB/continuum-electrostatics calculations or alchemically transformed particles. In benchmarks across four charged protein-ligand systems, EID achieved improved predictive accuracy and more consistent cross-system performance than both comparison methods. Because EID operates as a postprocessing step requiring no additional simulations or PB calculations, it provides a rigorous, immediately deployable solution for charge-changing free energy calculations.

Runduo Liu, Wanyi Huang, Yufen Yao et al. · 0 citations
Jul 2026

Native Contact Ratio as a Topological Metric for Machine Learning Based Molecular Docking

Accurate identification of near-native ligand binding poses is a central challenge in structure-based drug design. From a physical point of view, the successful construction of a protein–ligand complex structure is dependent on whether protein and ligand can form enough atomically pairwise interactions that result in a global energy minimum. In this work, we report a machine learning scoring strategy for protein–ligand screening which explicitly considers the Native Contact Ratio (NCR), a topology inspired metric that quantifies the preservation of protein–ligand interfacial contacts as well as interaction energy. This physics-awared supervision strategy provides a simple but efficient gradient field that faithfully reflects the complicated protein energy landscape than conventional 3D coordinate-based objectives. Building on this principle, we present DeepNCR, an energy-informed Transformer framework that encodes approximate Coulombic and dispersive interaction potentials across the protein–ligand binding interface. Furthermore, we introduce a feature pruning step that compresses the interaction tensor from 1470 to 868 dimensions, further improving signal-to-noise ratio and directing model attention toward the interaction motifs critical for binding specificity. The model optimizes topological objectives and at inference drives pose refinement through a differentiable hybrid gradient field integrating predicted NCR and AutoDock Vina energetics. Extensive evaluation on the CASF-2016 benchmark and the 3D-DISCO cross-docking data set demonstrates consistently high performance: a Top-1 docking success rate of 94.7%, a 1% Enrichment Factor of 21.21 in virtual screening, and a Top-1 cross-docking success rate of 34.8%. Mechanistic analysis reveals that NCR-guided optimization enables decoy escaping from local energy minima and drives the recovery of disrupted native interactions, confirming that NCR captures the physical determinants of binding rather than mere geometric proximity.

Zhenqiang Zhang, Zhihao Wang, Yang Liu et al. · 0 citations
Preprint Aug 2026

Accurate and Transferable Intermolecular Potential Based on Machine-Learned Molecular Electron Density

Machine-learned force fields (MLFFs) contain many learnable parameters and therefore require large training datasets. This poses a challenge for developing highly accurate, general-purpose MLFFs because generating high-quality ab initio reference data is computationally expensive. Classical empirical potentials offer a potentially inexpensive source of synthetic training data, but existing models often lack the accuracy needed to provide useful reference energies. Here, we introduce the density-based intermolecular potential (DensIP), a physics-based model of intermolecular interactions that uses machine-learned electron densities and only four universal parameters. We train and test DensIP on CCSD(T)/CBS interaction energies from DES15K, a dataset of dimers of small organic molecules. DensIP achieves sub-kcal/mol errors for dimers containing molecules absent from the training set, including molecules in non-equilibrium conformations, demonstrating strong transferability. We further show that DensIP can be applied to molecules as large as drug ligands. Notably, DensIP outperforms state-of-the-art general-purpose MLFFs for long-range interactions, making it a promising approach for generating accurate synthetic training data at scale.

Dahvyd Wing, Mihail Bogojeski, Szabolcs Góger et al. · 0 citations
Open access Jul 2026

Dual-Coordinate Relative Free Energy Simulations Using Machine-Learned Interatomic Potentials

The integration of machine-learned interatomic potentials (MLIPs) into free energy simulations (FES) offers the promise of near-quantum mechanical accuracy at a reduced computational cost. However, employing MLIPs for alchemical relative free energy calculations remains challenging since MLIPs are typically not trained to handle the unphysical intermediate states required for such transformations. In previous work, we demonstrated how atoms and molecules can be gradually decoupled in systems fully described by MLIPs by manipulating the neighbor list and introducing an artificial offset to interatomic distances. Here, we extend this approach to general alchemical transformations between two states and apply the methodology to computing relative solvation free energies (RSFEs) between arbitrary solute pairs in systems fully described by an unmodified MLIP. We employ a dual-coordinate approach where both solutes are explicitly present but do not interact with one another. By introducing λ-dependent distance offsets to the neighbor list, we perform a simultaneous transformation: smoothly decoupling the first solute from the solvent while coupling the second solute to the same environment. To enforce spatial overlap without modifying the internal intramolecular dynamics of either solute, we utilize a harmonic ″anchor″ restraint to a central atom on each molecule. We demonstrate the robustness of this method using the MACE-OFF23(S) potential across a diverse set of small molecule pairs, including transformations between species with no chemical similarity (e.g., toluene to tetrahydrofuran). The method’s accuracy is validated through thermodynamic cycle closure by combining direct RSFE calculations with absolute solvation free energy (ASFE) calculations. The methodology achieves excellent internal consistency, with cycle closure errors less than or equal to ±0.11 kcal/mol, well within the estimated statistical uncertainty. By requiring only a single energy evaluation per simulation step and avoiding architecture-specific modifications and retraining, this anchor-based dual-coordinate approach provides a highly adaptable and efficient route for relative alchemical FES with modern MLIPs.

Anna Katharina Picha, S. Boresch · 0 citations