Skip to content
Open access

Cross-docking and redocking reveal distinct determinants of success in physics-based and AI-driven binding pose prediction in protein–ligand complexes

Sep 2026 · RSC Advances · 0 citations · 25 references
Medicine

Abstract

Protein–ligand pose prediction is central to structure-based drug discovery, yet the relative performance of physics-based and AI-driven methods under realistic cross-docking conditions remains insufficiently characterized. Here, we compare physics-based docking methods (AutoDock4, AutoDock Vina, and DOCK 6) with data-driven approaches, including the deep-learning model GNINA 1.3 and the diffusion-based frameworks AlphaFold 3, Boltz-2, and DiffDock. Performance was evaluated using standardised redocking and cross-docking protocols across three Alzheimer's disease targets representing distinct binding-site architectures: acetylcholinesterase (AChE; deep gorge), β-secretase 1 (BACE1; flexible flap-controlled site), and glycogen synthase kinase-3β (GSK-3β; open, solvent-exposed pocket). Physics-based methods were competitive during redocking but showed substantial performance reductions under cross-docking, whereas diffusion-based approaches generally maintained higher cross-docking accuracy. GNINA 1.3 rigid achieved an 87.7% minimum heavy-atom RMSD success rate during redocking, which decreased to 13.5% during cross-docking, whereas AlphaFold 3, Boltz-2, and DiffDock achieved cross-docking success rates of 93.1%, 89.6%, and 85.7%, respectively. AlphaFold 3 consistently outperformed Boltz-2 despite its smaller training set, suggesting that predictive performance is influenced not only by training-data volume but also by factors such as model architecture and confidence calibration. Training-overlap analysis further showed that AI-based methods retained substantial failure rates even for complexes represented in their training data, indicating that training-data overlap alone does not ensure reliable pose prediction. Under the current protocol conditions, rigid docking outperformed flexible protocols, while flexible-docking pocket volumes showed more restricted sampling relative to experimental holo structures. Among the GNINA 1.3 configurations, CNN rescoring with refinement produced the highest pose-recovery success rates, followed by CNN rescoring alone and the default Vina/empirical scoring approach in cross-docking. Receptor conformational preference was target-dependent: holo structures provided higher docking accuracy for AChE and BACE1, whose ligand-bound cavities exhibited greater structural complexity and geometric confinement that favoured pose discrimination, whereas the apo GSK-3β structure contained a larger, more solvent-exposed cavity that improved ligand accessibility and docking performance. Overall, these findings demonstrate the importance of cross-docking and training-overlap-aware evaluation for assessing docking performance under realistic conditions and provide cavity-topology-based considerations for selecting docking strategies in structure-based drug discovery.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.