SurroDock: A Deep Learning Surrogate for Accelerated Pre-Docking Ligand Prioritization in Structure-Based Virtual Screening
The rapid expansion of make-on-demand and public chemical libraries has made exhaustive docking-based structure-based virtual screening increasingly difficult. This study introduces SurroDock, a lightweight deep-learning surrogate designed to approximate AutoDock Vina docking scores from low-cost two-dimensional molecular features, serving as a practical pre-filter for docking. SurroDock was evaluated for estrogen receptor alpha using two distinct conformations: an agonist-bound (PDB ID: 1GWR) and an antagonist/SERM-bound (PDB ID: 3ERT). The dataset comprised approximately 334,000 unique compounds curated from the NCI Open Database, PubChem, and BindingDB, all docked using a standardized AutoDock Vina workflow. The model was trained on concatenated 2D molecular representations comprising Morgan fingerprints, MACCS keys, RDKit physicochemical descriptors, Vina-inspired ligand descriptors, atom-pair fingerprints, and 2D pharmacophore fingerprints. The docking-score distributions differed substantially between receptor states, with 3ERT exhibiting more favorable scores than 1GWR and weak inter-state score correlation supporting state-specific modeling. Using the integrated Unified-200k training set (200,000 compounds randomly sampled per receptor from the three docked sources), SurroDock achieved strong held-out validation performance, with R2 values of approximately 0.88 for 1GWR and 0.93 for 3ERT. In retrospective screening-style evaluation, SurroDock recovered substantial fractions of Vina’s top-ranked compounds at the top-1% recall (Recall@1%) of approximately 0.57 and 0.61 for 1GWR and 3ERT, respectively, yielding corresponding enrichment factors (EF@1%) of approximately 57-fold and 61-fold relative to random selection. Overall, the results indicate that 2D-based docking-score surrogate modeling can provide a reproducible and retrainable strategy for large-scale structure-based virtual screening by concentrating docking resources on a smaller, enriched subset of compounds. Because SurroDock emulates a docking scoring function rather than experimental binding affinity, its predictions should be used as prioritization aids and complemented by confirmatory docking, pose inspection, and experimental validation.