Ultralarge virtual screenings (ULVSs) evaluate billions of molecules for drug discovery but face cost, flexibility and scalability limits. We introduce AdaptiveFlow, an open-source platform that makes ULVSs more accessible, scalable and efficient and supports artificial intelligence (AI) and machine learning (ML) method development. AdaptiveFlow provides a screening-ready version of the Enamine REAL Space, to our knowledge the largest library of ready-to-dock, drug-like molecules, comprising 69 billion compounds, also available in SELFIES format. An 18-dimensional grid of molecular properties prioritizes promising chemical subspaces, with optional active learning, reducing computational costs by orders of magnitude. AdaptiveFlow integrates >1,500 docking protocols, including GPU-accelerated and ML-based methods, and achieves near-linear scaling on up to 5.6 million CPUs in the Amazon Web Services cloud. We identified nanomolar inhibitors of two disease-relevant targets, ferroptosis suppressor protein 1 (FSP1) and poly(ADP-ribose) polymerase 1. Co-crystal structures provided mechanistic insights into FSP1 inhibition. AdaptiveFlow enables drug discovery at unprecedented scale and supports the development of AI-driven methods.
Domiziana Cecchini, AkshatKumar Nigam, Ming Tang et al.· Nature Biotechnology· 0 citations
Computational screening of giga-scale chemical spaces opens a cost-effective path to high-quality hit identification, providing entry points for drug discovery. As these on-demand spaces grow and successful applications multiply, rigorous blind benchmarks like CACHE Challenges provide important performance metrics for computational tools. Here, we report the first application of the V-SYNTHES2 synthon-based screening approach to the 173-billion-compound Enamine xREAL Space, a 16-fold expansion beyond its previous benchmarks, demonstrating near-linear computational scaling with only a 10–15% increase in cost relative to the 11-billion-compound REAL Space. We applied this workflow in CACHE Challenge #2, targeting the RNA-binding site of NSP13 (SARS-CoV-2), and CACHE Challenge #4, targeting the tyrosine kinase-binding domain of CBLB, both pockets lacking established pharmacology and representing extreme hit-finding challenges. Under blinded, independently validated conditions, V-SYNTHES2 ranked among the top-performing submissions: the 8% hit rate for NSP13 exceeded the field average of 2.3% and placed the approach among the top three workflows, while for CBLB, one compound meeting predefined hit criteria was identified. These results demonstrate that V-SYNTHES2 maintains robust performance at giga-scale on ligand-depleted targets, precisely the conditions where data-driven approaches would face fundamental limitations, and establish a quantitative performance baseline for synthon-based screening of hundred-billion-compound chemical spaces.
Mykola V. Protopopov, Olha Semenenko, Maryna Vasylchuk et al.· npj Drug Discovery· 0 citations
While make-on-demand libraries now span trillions of molecules, full library docking struggles beyond a few billion, motivating prioritization that recovers top-scoring compounds while evaluating only a fraction of a library. Here we introduce a similarity-based prioritization approach, ChemSTEP, and define the effective size of a library treated by any prioritization algorithm, Neff. ChemSTEP docks a representative seed set, selects diverse high-scoring “beacons”, and iteratively traverses the library through cycles of beacon selection, similarity search, and docking. Retrospectively on eight targets, ChemSTEP recovered over 75% of high-scoring compounds while docking less than 5% of a library. We then tested ChemSTEP prospectively against AmpC β-lactamase using a 13.2 billion molecule library. Because AmpC recognizes negatively charged inhibitors, we explicitly docked all 360 million library anions, synthesizing and testing 241 high-ranking ones in parallel to the ChemSTEP 13.2B run. Compared with previous docking of 99 million and 1.7 billion molecules against AmpC, the 13.2 billion library had higher hit-rates (2% vs 25% vs 37%, respectively) and found more potent compounds. Meanwhile, ChemSTEP retrieved 80% of the 241 high-ranking compounds within the first 0.5% docked. Trillion-molecule libraries might be in reach with this approach.
Olivier Mailhot, Katie L. Holland, Lu Paris et al.· ACS Central Science· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.