Jul 2026· Journal of the American Chemical Society· Vol 148, pp. 29908 - 29920· 0 citations· 70 references
Medicine
TL;DR
ComBINAUT is an automated parallel synthesis platform that generates diverse chemical scaffolds to accelerate hit validation and refinement, enabling the efficient exploration of chemical space and the rapid discovery of novel ligands.
Abstract
Virtual screening (VS) is a powerful approach to exploring a vast chemical space, encompassing libraries of millions to billions of compounds. However, the low hit rates of VS require testing numerous candidates to validate true binders, followed by iterative optimization cycles, which makes experimental validation costly and time-consuming. Here, we report COMBINAUT, an automated parallel synthesis platform that generates diverse chemical scaffolds to accelerate hit validation and refinement. Using a faculty-wide collection of in-house building blocks, the system enables enumeration of over 22.9 million compounds, each designed for parallelized synthesis within 32 h using repurposed solid-phase peptide synthesis equipment. Using this platform, we performed large-scale VS targeting the allosteric pocket of the immuno-oncology target, C–C chemokine receptor 2 (CCR2). Our approach facilitated the rapid synthesis and testing of 100 VS hits spanning diverse molecular architectures. In radioligand binding assays, we successfully validated nine hits with distinct scaffolds, including completely novel CCR2 ligand chemotypes. Iterative hit-to-lead optimization using the automated workflow produced cell-active CCR2 antagonists. This work demonstrates the synergy of automated synthesis and VS, enabling the efficient exploration of chemical space and the rapid discovery of novel ligands.
Computational screening of giga-scale chemical spaces opens a cost-effective path to high-quality hit identification, providing entry points for drug discovery. As these on-demand spaces grow and successful applications multiply, rigorous blind benchmarks like CACHE Challenges provide important performance metrics for computational tools. Here, we report the first application of the V-SYNTHES2 synthon-based screening approach to the 173-billion-compound Enamine xREAL Space, a 16-fold expansion beyond its previous benchmarks, demonstrating near-linear computational scaling with only a 10–15% increase in cost relative to the 11-billion-compound REAL Space. We applied this workflow in CACHE Challenge #2, targeting the RNA-binding site of NSP13 (SARS-CoV-2), and CACHE Challenge #4, targeting the tyrosine kinase-binding domain of CBLB, both pockets lacking established pharmacology and representing extreme hit-finding challenges. Under blinded, independently validated conditions, V-SYNTHES2 ranked among the top-performing submissions: the 8% hit rate for NSP13 exceeded the field average of 2.3% and placed the approach among the top three workflows, while for CBLB, one compound meeting predefined hit criteria was identified. These results demonstrate that V-SYNTHES2 maintains robust performance at giga-scale on ligand-depleted targets, precisely the conditions where data-driven approaches would face fundamental limitations, and establish a quantitative performance baseline for synthon-based screening of hundred-billion-compound chemical spaces.
Mykola V. Protopopov, Olha Semenenko, Maryna Vasylchuk et al.· npj Drug Discovery· 0 citations
Ultralarge virtual screenings (ULVSs) evaluate billions of molecules for drug discovery but face cost, flexibility and scalability limits. We introduce AdaptiveFlow, an open-source platform that makes ULVSs more accessible, scalable and efficient and supports artificial intelligence (AI) and machine learning (ML) method development. AdaptiveFlow provides a screening-ready version of the Enamine REAL Space, to our knowledge the largest library of ready-to-dock, drug-like molecules, comprising 69 billion compounds, also available in SELFIES format. An 18-dimensional grid of molecular properties prioritizes promising chemical subspaces, with optional active learning, reducing computational costs by orders of magnitude. AdaptiveFlow integrates >1,500 docking protocols, including GPU-accelerated and ML-based methods, and achieves near-linear scaling on up to 5.6 million CPUs in the Amazon Web Services cloud. We identified nanomolar inhibitors of two disease-relevant targets, ferroptosis suppressor protein 1 (FSP1) and poly(ADP-ribose) polymerase 1. Co-crystal structures provided mechanistic insights into FSP1 inhibition. AdaptiveFlow enables drug discovery at unprecedented scale and supports the development of AI-driven methods.
Domiziana Cecchini, AkshatKumar Nigam, Ming Tang et al.· Nature Biotechnology· 0 citations
It is shown that prospective structure selection, rather than structure generation, represents the primary bottleneck in ensemble-based VS, highlighting an urgent need for novel structural descriptors to identify high-performing conformations.
Jaeoh Shin, K. Joo, Jejoong Yoo· Journal of Chemical Informat...· 0 citations
The principal constraint on early-stage medicinal chemistry in an academic setting is rarely the supply of chemical ideas, but the cycle time required to convert them into tested compounds. Here we benchmark an integrated platform that couples ensemble virtual screening, automated continuous-flow library synthesis with inline scavenging, and biological profiling. The platform was evaluated deliberately on a target-scaffold combination for which the pharmacology is already established: the N-benzylpiperidine carboxamide class, identified in our earlier virtual screening campaign against acetylcholinesterase (AChE)[1] and structurally anchored to the approved therapeutic donepezil (1). This choice makes platform performance, the measured variable. An 84-member virtual library was designed, triaged by ensemble docking, and synthesized on an automated flow platform; 54 members were isolated in ≥95% purity. Single-point screening at 5 μM identified ten compounds with ≥70% AChE inhibition, and dose-response determination against electric eel AChE (eeAChE) gave three sub-200 nM inhibitors: 69 (91 nM), 8 (93 nM) and 12 (117 nM), each exceeding galantamine (237 nM) and approaching donepezil (1, 42 nM) under identical assay conditions. Cytotoxicity against VERO cells was uniformly low (IC50 > 250 μM), giving selectivity indices above 2000. Molecular dynamics simulations and twelve single-crystal X-ray structures provide a structural basis for the observed structure-activity relationships, identifying a C-Br···O halogen bond to Asp72 as the origin of the ortho-bromo preference, and a binding-mode reversal that accounts for the loss of potency on benzylpiperazine extension. The complete cycle was executed in under three months by a three-person team. We report this as a measured platform capability rather than as accelerated drug discovery allowing the rapid generation of first-round leads.
A. Reinhardt, A. Theron, Asongwe Lionel Ateh Tantoh et al.· European journal of medicina...· 0 citations
This work systematically evaluates the performance of a popular ML-based docking method, DiffDock-Pocket, on high-throughput screening (HTS) data sets derived from the PubChem BioAssay database, a premier source of bioactivity data.
Furyal Ahmed, M. Soellner, Charles L. Brooks· Journal of Chemical Informat...· 0 citations
Virtual screening (VS) on small molecules aims to identify promising drug candidates against protein targets from expansive chemical libraries by balancing the core requirements of accurate scoring and efficient search against the inherent trade‐off between accuracy and speed. This survey provides a comprehensive review of how Artificial Intelligence and Machine Learning (AI/ML) are redefining this landscape across three critical dimensions. First, we examine the evolution of AI‐driven scoring functions, which utilize AI/ML models to capture complex structure–activity relationships from massive biochemical datasets, significantly enhancing structure‐ and ligand‐based evaluations beyond traditional heuristics. Second, we summarize the emergence of efficient search algorithms that iteratively prioritize informative compounds to reduce search efforts by orders of magnitude. Third, we review the paradigm shift toward generative molecular design, making VS transition from screening fixed libraries to the
de novo
generation of molecules optimized for specific structural contexts and multi‐objective properties. This review outlines the transition toward end‐to‐end, adaptive discovery systems that ensure computational hits are biologically potent, structurally optimized, and synthetically accessible.
Yifei Wang, Nupur Bansal, Shiyun Wa et al.· WIREs Computational Molecula...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.