Skip to content
Open access

Prioritising search for virtual screening via preliminary interpretable low-feature likelihood-based rankings of drug-target activity measures.

Jul 2026 · BMC Bioinformatics · 0 citations
Medicine

TL;DR

This article proposes an offline/online method that Wisely selects a small number of easy to compute features of both the amino acid sequence of the proteins and the molecular structure of the ligands, and discretises their domains, which induces a low-dimensional finitisation of proteins' and ligands' chemical spaces.

Abstract

Background

Current AI-based Virtual Screening (VS) methods seek to manage ultra-large molecular libraries. To this end, they develop increasingly efficient heuristics to rank ligands by their predicted activity against a target protein. However, these methods remain computationally demanding due to the billion-scale compound libraries that must be evaluated without prior, informed guidance.

Results

This article proposes an offline/online method that: (1) Wisely selects (once and forall, offline phase) a small number of easy to compute features [Formula: see text] of both the amino acid sequence of the proteins ([Formula: see text]) and the molecular structure of the ligands ([Formula: see text]), and discretises their domains; this induces a low-dimensional finitisation of proteins' and ligands' chemical spaces. (2) Given a target protein [Formula: see text], immediately returns (online phase) a likelihood-based ranking of the classes of the ligands' chemical space, in descending order of the estimated probability that molecules in each class will achieve a satisfactory activity measurement against [Formula: see text]. This enables any VS method to prioritise the search to the most promising subsets of candidates. To ensure statistically robustness, our offline feature selection: (a) leverages knowledge stemming from a huge dataset of 2 559 403 entries (ligand-protein activity measurements) obtained by unifying the most representative sources regarding biochemical kinetics (Brenda, Sabio-rk, BindingDB) and augmented with 3781 features computed by 7 well-known third-party software tools; (b) explicitly aims at low-dimensional coarse-domain feature spaces; (c) takes proper countermeasures to prevent biases in the source data and overfitting; (d) supports iterative improvement of [Formula: see text] via an anytime offline algorithm and means to interactively exclude features deemed uninformative upon rankings inspection; (e) supports intepretability of the rankings by enabling inspection of the features' values characterising each ligand class.

Conclusions

By evaluating our rankings on evaluation data (from PDBbind and additional BindingDB entries unsuitable for accurate statistical analysis), we demonstrate their effectiveness for library prioritisation. Specifically, our findings indicate that approximately 60% of the high-affinity ligands occur in the top 25% ranked ligands' classes, while 85% fall within the top 50%. Furthermore, we conduct retrospective analysises using AutoDock Vina scores for over 260 000 molecules across 58 medically relevant targets. Results demonstrate that our method cuts the number of dockings needed to retrieve an equivalent set of hits by up to [Formula: see text] on average versus unguided screening.

Read PDF

Similar papers

Jul 2026

Real-World Assessment of Machine-Learned Docking Using Bioassay-Derived Benchmarks

This work systematically evaluates the performance of a popular ML-based docking method, DiffDock-Pocket, on high-throughput screening (HTS) data sets derived from the PubChem BioAssay database, a premier source of bioactivity data.

Furyal Ahmed, M. Soellner, Charles L. Brooks · 0 citations
#generative ai Review Sep 2026

Machine Learning‐Aided Small‐Molecule Virtual Screening: Recent Advances and Future Perspectives

Virtual screening (VS) on small molecules aims to identify promising drug candidates against protein targets from expansive chemical libraries by balancing the core requirements of accurate scoring and efficient search against the inherent trade‐off between accuracy and speed. This survey provides a comprehensive review of how Artificial Intelligence and Machine Learning (AI/ML) are redefining this landscape across three critical dimensions. First, we examine the evolution of AI‐driven scoring functions, which utilize AI/ML models to capture complex structure–activity relationships from massive biochemical datasets, significantly enhancing structure‐ and ligand‐based evaluations beyond traditional heuristics. Second, we summarize the emergence of efficient search algorithms that iteratively prioritize informative compounds to reduce search efforts by orders of magnitude. Third, we review the paradigm shift toward generative molecular design, making VS transition from screening fixed libraries to the de novo generation of molecules optimized for specific structural contexts and multi‐objective properties. This review outlines the transition toward end‐to‐end, adaptive discovery systems that ensure computational hits are biologically potent, structurally optimized, and synthetically accessible.

Yifei Wang, Nupur Bansal, Shiyun Wa et al. · 0 citations
Open access Sep 2026

Data driven selection of consensus docking pipelines for structure based hit identification

Structure-based virtual screening (SBVS) is a cornerstone of computer-aided drug design, yet its success depends on selecting a combination of docking tools, scoring function (SF), and ranking strategies. MolDockLab addresses this challenge with an automated, data-driven framework that optimizes SBVS workflows for a protein target, balancing predictive performance and computational efficiency. It systematically explores combinations of five docking engines, 15 SF, and three consensus ranking strategies using a calibration set of ≈ 200 compounds with known bioactivity, and applies the best-correlating workflow to the larger screening library. Final hit selection from the top 1% integrates protein-ligand interaction profiler (PLIP)-derived interaction fingerprints, structural-diversity assessment, and expert visual inspection. In a retrospective evaluation on the epidermal growth factor receptor (EGFR), the chosen pipeline achieved a Spearman correlation of 0.36 and an enrichment factor (EF) at 10% of 1.57, consistent with calibration. Prospectively, for the energy coupling factor transporters (ECF-T)-a challenging transmembrane target with a cryptic binding site and no co-crystallized ligand-the pipeline reached a correlation of 0.45 and enrichment of 3.13. Post-processing enabled in vitro confirmation of two chemically novel inhibitors rivaling the most potent ECF-T inhibitors reported to date.

Hamza Agha, Y. Ibrahim, Michael Backenköhler et al. · 0 citations
Jul 2026

Abstract A024: FastBindRank, a novel, scalable method for high-fidelity virtual screening of ultra-large chemical libraries idendtifies novel HDAC11 inhibitors

FastBindRank is presented, a distillation-based framework that transfers the predictive power of a high-accuracy structure-based model (Boltz-2) into a computationally efficient sequence-based surrogate and can identify functionally active novel compounds from ultra-large virtual screening.

J. Dai, Yueyue Wang, N. Shan et al. · 0 citations
Preprint Aug 2026

PETA:Parameter-Efficient Test-Time Adaptation for Virtual Screening

This work forms the specialization of pretrained virtual screening models to individual pockets as a test-time adaptation problem and proposes PETA, a parameter-efficient framework that directly adapts pretrained model at test time and outperforms both pretrained and fully retrained baselines while updating only the LayerNorm parameters.

Jia-Qi Lin, Yinghua Yao, Changran Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.