Skip to content
Open access

Evaluating molecular docking for binding affinity predictions: a systematic analysis of key parameters and the utility of AlphaFold2 structures for the Schrödinger dataset

Aug 2026 · Journal of Computer-Aided Molecular Design · Vol 40 · 0 citations · 60 references
Medicine

TL;DR

Docking should be considered as an important and computationally inexpensive reference baseline for binding affinity prediction, and the scoring function and the protein structure are the most important factors for binding affinity accuracy in rigid docking with the MOE software.

Abstract

Molecular docking is one of the most established methods in computational drug discovery, due to its balance of speed and accuracy. However, the accuracy of docking results depends on a number of different parameters, and systematic reference data for comparisons to more advanced methods for binding affinity prediction are still scarce. This study assesses the impact of key parameters on the accuracy of binding free energy estimates from docking, using nine benchmark systems with 278 high-affinity ligands. Using the Molecular Operating Environment (MOE), we evaluated combinations of three receptor structures (two crystal structures, one AlphaFold2 model), two force fields, two scoring functions, two receptor flexibility settings, and two statistical evaluation schemes. The performance of the docking approaches is measured based on the squared Pearson’s correlation coefficient (R²), the root mean square error (RMSE) with respect to the experimental binding affinities, as well as the mean signed error (MSE) and Kendall’s tau for individual targets and the full dataset. The results show that the scoring function and the protein structure are the most important factors for binding affinity accuracy in rigid docking with the MOE software. Amber10:EHT and MMFF94x force fields had the same average Rmean2 value, but Amber10:EHT had a lower average RMSEmean. AlphaFold2 protein models yielded lower binding affinity accuracy and higher errors compared to experimental crystal structures, although induced fit docking improved results. Using the original benchmark, we also compared several docking programs. DOCK6 and MOE performed best, with mean R² values of about 0.49 and 0.40, respectively. The remaining docking programs did not outperform a molecular weight regression baseline. For a subset of four targets (CDK2, JNK1, P38, TYK2) evaluated in previous work, the performance of the optimized DOCK6 and MOE protocols produced correlation coefficients similar to those reported for certain MM/PBSA, FMO, and Boltz2 implementations evaluated on the same target subset. This raises questions about potential dataset biases, the structural preparation, or the implementation of those methods. Docking therefore should be considered as an important and computationally inexpensive reference baseline for binding affinity prediction.

Read PDF

Similar papers

Open access Jul 2026

Benchmarking Docking Protocols on Predicting Alternative Binding Modes

Assessment of pose prediction methods when the bound structure of a reference ligand is known and the likely binding mode(s) of a related compound are needed, and this work focuses on cases where the new compound has multiple potential binding modes.

Ažbeta Kubincová, S. S. Çınaroğlu, Jianna Ongsioco et al. · 0 citations
Sep 2026

Performances and Critical Limitations of Docking and Protein–Ligand Interactions for Efficacy-Driven Virtual Screening Targeting Class A GPCRs

Structure-based virtual screening of chemical libraries is an established and widely used strategy for identifying novel ligands for G-protein-coupled receptors. An enhancement based on integrating protein–ligand interaction with docking has previously been proposed, but its actual impact on improving screening outcomes has remained unclear. Here, we present a comprehensive assessment based on systematic benchmarking using a diverse set of class A G-protein-coupled receptors and different approaches to represent protein–ligand, including dynamic patterns extracted from molecular dynamics simulations. Our results demonstrate that the combined approach overall improves the efficacy bias of selected ligands as compared to docking alone (ranking by scoring function). All tested variations prove broadly functional; however, the most sophisticated one─incorporating simulation and a learning model─emerges as the most robust alternative for a prospective setting. The analysis of two prospective cases, the design of both agonists and antagonists of CNR1 and the more challenging search for CXCR4 nonpeptidic agonist, reveals both great potential and inherent structural limitations, highlighting the need for an accurate and suitable three-dimensional structure.

Luca Chiesa, G. Bret, Severine Schneider et al. · 0 citations
Open access Aug 2026

Probe-Atom Distributions Obtained from Mixed-Solvent Molecular Dynamics Improve the Scoring of Docking Calculations

Structure-based virtual screening (VS) is widely used for the computational selection of drug candidates from compound libraries. Protein–ligand docking calculations are often performed as key steps in the early stages of this process. However, current docking calculations have limited accuracy. Thus, improvements are needed to more efficiently identify promising drug candidates. In this study, we performed mixed-solvent molecular dynamics (MSMD) simulations using four types of probe molecules to improve the accuracy of large-scale VS. We proposed a method for the modification of the docking scoring function for five selected atom classifications (XS_types). This approach integrated the grid free energy derived from the relevant atoms across the probe molecules. VS experiments conducted on nine target proteins showed improved accuracy, with the average EF1% increasing from 6.65 to 7.36. Our method may facilitate drug discovery with higher accuracy than that of conventional methods.

Unknown authors · 0 citations
Open access Jul 2026

Reliability of AI Methods in Drug Discovery: Evaluation of Boltz‑2 for Structure and Binding Affinity Prediction

An extensive evaluation of Boltz-2 using two large-scale data sets shows that Boltz-2 lacks the energetic resolution required for lead identification, highlighting the necessity of employing physics-based methods for the reliability and refinement of AI-derived models.

S. Wan, Xibei Zhang, Xiao Xue et al. · 3 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.