Similar papers
Real-World Assessment of Machine-Learned Docking Using Bioassay-Derived Benchmarks.
The rapid expansion of compound libraries has significantly advanced drug discovery, especially through ultralarge library screenings that provide access to vast chemical spaces. However, the sheer scale of these libraries introduces substantial challenges in the early stages of drug discovery. While it is true that searching larger libraries can improve the hit rate of a virtual screening campaign, as the available chemical space has increased to over 64 billion molecules, identifying relevant information becomes a complex and time-consuming task. Machine learning (ML)-based docking methods and scoring functions offer a potential solution by providing increased speed and scalability. However, much of their reported success is based on benchmark data sets that rely on constructed decoys and ligand sets, which often have hidden biases that can artificially inflate performance and fail to capture the challenges of real-world screening. In this work, we systematically evaluate the performance of a popular ML-based docking method, DiffDock-Pocket, on high-throughput screening (HTS) data sets derived from the PubChem BioAssay database, a premier source of bioactivity data. By using HTS data sets as a more realistic benchmark, we aim to provide a clearer picture of how ML models perform in practical virtual screening scenarios compared to traditional physics-based docking approaches. Our work highlights the strengths and limitations of current ML methods and offers insights into their reliability and applicability to prospective applications in drug discovery.
Modeling the Sensitivity of Large-Scale Virtual Screening to Scoring Function Accuracy, Artifacts, and Library Composition
Large library docking has emerged as a productive approach for ligand discovery, yet a quantitative framework for understanding how docking performance responds to methodological improvements has been lacking. Here, we develop such a framework by modeling large-scale experiments from three previously published docking campaigns, in which 2,682 ligands had been synthesized and tested across the scoring landscape (poor scores, mediocre scores, high scores). The observed experimental hit-rate curves can be reproduced by a simple bivariate normal distribution model, where docking score is interpreted as a noisy predictor of binding free energy. To account for the plateauing and subsequent drop in hit rates often seen at highly favorable docking scores, we add a term for high-ranking docking artifacts, a phenomenon we observe across targets. From this model, three predictions about the sensitivity of docking performance emerge. First, even slight improvements in scoring accuracy would substantially improve both hit rates and hit affinities: quantitatively, a 0.1 increase in the correlation between docking score and binding affinity would justify accepting a ∼10-fold increase in computational cost per molecule, arguing for reinvestment in scoring function accuracy in library docking. Second, docking artifacts, while hard to anticipate, can come to dominate top-scoring lists as libraries grow. Physically testing molecules across a range of log-normalized ranks (pProp) is therefore essential to identify the peak hit rate for a given campaign. Third, prefiltering a library to enrich for molecules with appropriate physicochemical features increases the intrinsic hit rate and substantially boosts docking performance, particularly at tera-scale, with effects comparable to a meaningful improvement in scoring accuracy. Beyond docking, the model’s parameters (affinity distribution, score-affinity correlation, artifact frequency) can be fit to any screening method with sufficient experimental data, providing an objective basis for benchmarking and comparing virtual screening approaches. These findings offer a practical framework for optimizing large-scale virtual screening as chemical libraries continue to grow.
Topological deep learning for drug–target interaction, virtual screening, and docking scoring: a practical, benchmark-driven review
A decision-oriented taxonomy and a benchmark-driven evaluation playbook that specifies minimum standards for splits, metrics, baselines, and ablations to isolate the topological contribution are presented.
PETA:Parameter-Efficient Test-Time Adaptation for Virtual Screening
Accurately ranking active ligands for a target protein pocket from massive chemical libraries remains a central challenge in virtual screening. DrugCLIP and its recent extensions substantially accelerate this process by encoding protein pockets and molecules into a shared embedding space. Despite this progress, further performance improvements typically require retraining the entire model, incurring substantial computational overhead and making target-specific customization inefficient. In this work, we formulate the specialization of pretrained virtual screening models to individual pockets as a test-time adaptation problem and propose PETA, a parameter-efficient framework that directly adapts pretrained model at test time. Given a target pocket, PETA constructs pocket-specific negatives through molecular diffusion and chemical validity filtering, and further moves them toward the reference ligand retrieved from structural databases via embedding-space mixup to create more challenging ranking tasks. A ranking objective then places greater emphasis on suppressing high-scoring invalid candidates that could contaminate the top-ranked screening results, providing structured supervision for lightweight adaptation. Experiments across diverse benchmarks demonstrate that this lightweight, pocket-specific adaptation outperforms both pretrained and fully retrained baselines while updating only the LayerNorm parameters, which account for approximately $0.03\%$ of the full model.
BoltzMol-1: Towards Reliable Virtual Screening for Fast and Cost-Effective Hit Discovery
We present BoltzMol-1, a small-molecule hit discovery pipeline, centered on an optimized version of Boltz-2, explicitly adapted for prospective discovery. Reliable hit discovery that generalizes across target classes (rather than only the well-characterized families that dominate existing ligand data) would broaden the range of biology accessible to small-molecule intervention and reduce reliance on resource-intensive high-throughput screening. Towards this goal, the system prioritizes compounds for rapid experimental validation by coupling model-driven ranking with streamlined procurement from commercial catalogs. To improve developability at the point of selection, we introduce a suite of ADMET models for kinetic solubility (logS), lipophilicity (logD), and Caco-2 permeability. These models act as an early triage layer, systematically filtering out compounds with unfavorable physicochemical and absorption properties prior to synthesis or purchase. Across a panel of ten targets (most with no representation in the underlying affinity training data) we observe strong prospective performance on challenging systems. Functional actives or binders were identified for 6 of 10 targets, despite modest experimental budgets of 28-96 compounds per target. These results include successes on receptors and enzymes traditionally considered difficult for structure- or ligand-based approaches. Collectively, this work establishes a practical framework for low-throughput, cost-constrained discovery campaigns capable of delivering chemically tractable binders with favorable property profiles. Figure 1: Overview of the prospective virtual-screening campaigns across all targets. For each target, the panel shows the predicted protein-ligand complex together with the number of compounds tested, the number of confirmed actives/binders, and the assays used for screening and follow-up.
Evaluating BioEmu-Generated Kinase Ensembles Reveals Structure Selection as the Virtual Screening Bottleneck
Virtual screening (VS) is an essential tool in drug discovery to prioritize potential drug candidates from vast chemical space. One key challenge limiting its performance is accounting for protein conformational flexibility. While ensemble docking methods have been developed to address this challenge by incorporating multiple protein conformations, these methods often rely on computationally intensive physics-based simulations to sample the relevant conformational space. Generative machine learning models offer a highly promising, scalable, and high-throughput alternative to overcome the limitations of these traditional approaches. We therefore investigate whether conformational ensembles generated by BioEmu, a recently developed generative model, can improve VS performance for kinase targets. Using the DUD-E benchmark data set and a validated AutoDock-GPU protocol, we generated and analyzed nearly 1300 structures across 26 kinases (approximately 50 structures each). BioEmu produces structurally diverse ensembles with substantial performance variation among individual structures. However, ensemble methods employing consensus or best-score selection fail to improve upon, and often degrade, VS performance compared to crystal structure baselines. To investigate the source of this limitation, we quantified the relationship between KinCoRe-based conformational state classification and screening performance. By calculating the coefficient of determination (R2) across the kinase subset, we found that the structural features governing VS performance differ substantially from those defining standard conformational states, with KinCoRe classifications leaving over 84% of performance variance unexplained. This critical gap demonstrates that structural diversity alone is insufficient to guarantee screening success. We show that prospective structure selection, rather than structure generation, represents the primary bottleneck in ensemble-based VS, highlighting an urgent need for novel structural descriptors to identify high-performing conformations.