Skip to content
Open access

Beyond the Score: Fixed-Budget Benchmarking of Virtual Screening Integration Strategies for Decision-Centric Drug Discovery

Aug 2026 · International Journal of Molecular Sciences · Vol 27 · 0 citations · 33 references
Medicine

TL;DR

Target-level results showed substantial variability in MCS-containing workflows and limited benefits from adding docking without target-specific optimization, and a validated ligand-based predictor or a simple two-method rank-fusion scheme provided the highest observed mean hit recovery without requiring elaborate integration.

Abstract

Virtual screening (VS) workflows often combine structure- and ligand-based methods; however, their value depends on the number of compounds that can be tested. We benchmarked 20 fixed-budget strategies derived from molecular docking (GNINA CNN score), maximum common substructure (MCS) similarity, and a calibrated machine-learning (ML)-QSAR classifier across five pharmacologically diverse targets. Individual methods, best-rank and worst-rank fusion, mean-rank consensus, and sequential funnels were evaluated at 1%, 5%, and 10% library fractions, with every strategy selecting the same number of compounds. ML-QSAR was the strongest standalone method, recovering 47.6%, 81.6%, and 84.4% of actives at the three cutoffs. At the 1% budget, ML-QSAR achieved the highest mean hit recovery (47.6% recall; 99.2% precision). At 5% and 10%, best-rank fusion of QSAR and MCS produced the highest mean recall (83.2% and 86.4%). Among the sequential workflows, QSAR → MCS achieved the highest 1% hit recovery (45.2 ± 3.3% recall), whereas docking-first funnels consistently underperformed under the default, non-optimized conditions evaluated in this study. Target-level results showed substantial variability in MCS-containing workflows and limited benefits from adding docking without target-specific optimization. Under matched assay budgets, a validated ligand-based predictor or a simple two-method rank-fusion scheme provided the highest observed mean hit recovery without requiring elaborate integration.

Read PDF

Similar papers

Open access Sep 2026

Data driven selection of consensus docking pipelines for structure based hit identification

Structure-based virtual screening (SBVS) is a cornerstone of computer-aided drug design, yet its success depends on selecting a combination of docking tools, scoring function (SF), and ranking strategies. MolDockLab addresses this challenge with an automated, data-driven framework that optimizes SBVS workflows for a protein target, balancing predictive performance and computational efficiency. It systematically explores combinations of five docking engines, 15 SF, and three consensus ranking strategies using a calibration set of ≈ 200 compounds with known bioactivity, and applies the best-correlating workflow to the larger screening library. Final hit selection from the top 1% integrates protein-ligand interaction profiler (PLIP)-derived interaction fingerprints, structural-diversity assessment, and expert visual inspection. In a retrospective evaluation on the epidermal growth factor receptor (EGFR), the chosen pipeline achieved a Spearman correlation of 0.36 and an enrichment factor (EF) at 10% of 1.57, consistent with calibration. Prospectively, for the energy coupling factor transporters (ECF-T)-a challenging transmembrane target with a cryptic binding site and no co-crystallized ligand-the pipeline reached a correlation of 0.45 and enrichment of 3.13. Post-processing enabled in vitro confirmation of two chemically novel inhibitors rivaling the most potent ECF-T inhibitors reported to date.

Hamza Agha, Y. Ibrahim, Michael Backenköhler et al. · 0 citations
Jul 2026

Real-World Assessment of Machine-Learned Docking Using Bioassay-Derived Benchmarks

This work systematically evaluates the performance of a popular ML-based docking method, DiffDock-Pocket, on high-throughput screening (HTS) data sets derived from the PubChem BioAssay database, a premier source of bioactivity data.

Furyal Ahmed, M. Soellner, Charles L. Brooks · 0 citations

Systematic Evaluation of Graph Neural Networks for Ligand-Based Virtual Screening on ChEMBL Datasets

The performance of target-specific, ligand-based virtual screening models is strongly influenced by dataset characteristics, including data availability, class imbalance, and evaluation strategies. In this work, we perform a systematic evaluation of graph neural networks (GNNs) using a ChEMBL-derived dataset spanning 5,368 targets and 1.59 million activity records, capturing the long-tailed distributions and target-specific imbalances commonly observed in pharmaceutical data. Through a systematic evaluation of multiple GNN architectures, we identify guidelines for model selection: while the Graph Isomorphism Network (GIN) consistently outperforms others on datasets with >100 samples (a mean ROC–AUC up to 0.94), simpler architectures are more robust under extreme data scarcity. Critically, our comparative analysis of splitting strategies reveals that random sampling yields artificially optimistic performance due to structural overlaps, whereas similarity-aware clustering exposes a substantial generalization gap (AUC drop > 0.3), cautioning against prevailing evaluation practices. We further demonstrate that multi-task learning serves as an effective remedy for small-target instability, providing significant and consistent performance gains. To underscore its translational value, we deploy this comprehensive framework in a virtual screening campaign against Mcl-1, yielding a chemically optimized lead, C4 (Ki = 0.58 μM), with verified cellular efficacy. Our findings highlight the importance of task-aware benchmark design and offer a practical strategy for reliable GNN application in drug discovery.

Haihan Liu, Jiaqi Lin, Ying Fan et al. · 0 citations
Open access Sep 2026

AI-enhanced adaptive virtual screening of large libraries for ligand discovery.

Ultralarge virtual screenings (ULVSs) evaluate billions of molecules for drug discovery but face cost, flexibility and scalability limits. We introduce AdaptiveFlow, an open-source platform that makes ULVSs more accessible, scalable and efficient and supports artificial intelligence (AI) and machine learning (ML) method development. AdaptiveFlow provides a screening-ready version of the Enamine REAL Space, to our knowledge the largest library of ready-to-dock, drug-like molecules, comprising 69 billion compounds, also available in SELFIES format. An 18-dimensional grid of molecular properties prioritizes promising chemical subspaces, with optional active learning, reducing computational costs by orders of magnitude. AdaptiveFlow integrates >1,500 docking protocols, including GPU-accelerated and ML-based methods, and achieves near-linear scaling on up to 5.6 million CPUs in the Amazon Web Services cloud. We identified nanomolar inhibitors of two disease-relevant targets, ferroptosis suppressor protein 1 (FSP1) and poly(ADP-ribose) polymerase 1. Co-crystal structures provided mechanistic insights into FSP1 inhibition. AdaptiveFlow enables drug discovery at unprecedented scale and supports the development of AI-driven methods.

Domiziana Cecchini, AkshatKumar Nigam, Ming Tang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.