Closed-loop generative selection has become a workhorse of computational drug discovery: a learned generative model proposes candidate molecules, a fitness oracle scores them, the best are kept, and the model is retrained on this elite set before the next round. Despite its wide use, the method has lacked a rigorous convergence theory, largely because retraining the model each round breaks the Markov property on which classical evolutionary-algorithm analysis relies. We develop a self-contained theory of convergence and expected running time for this class of algorithms. By recovering a Markov structure on an enlarged state space, we show that elitism makes the search absorbing, and we prove almost-sure convergence together with a runtime bound that decomposes the search into the time spent escaping each fitness level. We then analyse the role of the model's memory---how much of the past it is trained on. When learning improves steadily with more data, deeper memory never hurts; when it does not, an exit-time analysis pinpoints the optimal memory depth and shows that excess memory can actually slow convergence. The theory extends to multi-objective search and to noisy oracles: we quantify how many repeated evaluations certify progress under light-tailed noise, and how robust estimators restore guarantees under heavy tails. Recast in terms of oracle evaluations - the true bottleneck in drug design - the analysis yields a concrete, evaluation-minimal strategy. Areproducible study confirms the predictions, including the surprising cost of excess memory. We close with three open problems.
K. Fackeldey, Christof Schütte· arXiv.org· 0 citations
Ultralarge virtual screenings (ULVSs) evaluate billions of molecules for drug discovery but face cost, flexibility and scalability limits. We introduce AdaptiveFlow, an open-source platform that makes ULVSs more accessible, scalable and efficient and supports artificial intelligence (AI) and machine learning (ML) method development. AdaptiveFlow provides a screening-ready version of the Enamine REAL Space, to our knowledge the largest library of ready-to-dock, drug-like molecules, comprising 69 billion compounds, also available in SELFIES format. An 18-dimensional grid of molecular properties prioritizes promising chemical subspaces, with optional active learning, reducing computational costs by orders of magnitude. AdaptiveFlow integrates >1,500 docking protocols, including GPU-accelerated and ML-based methods, and achieves near-linear scaling on up to 5.6 million CPUs in the Amazon Web Services cloud. We identified nanomolar inhibitors of two disease-relevant targets, ferroptosis suppressor protein 1 (FSP1) and poly(ADP-ribose) polymerase 1. Co-crystal structures provided mechanistic insights into FSP1 inhibition. AdaptiveFlow enables drug discovery at unprecedented scale and supports the development of AI-driven methods.
Domiziana Cecchini, AkshatKumar Nigam, Ming Tang et al.· Nature Biotechnology· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.