These results demonstrate that calibrated AL strategies can overcome data acquisition bottlenecks and train generalizable property predictors able to extrapolate to OOD molecules.
Abstract
Machine learning (ML) has accelerated molecular discovery, yet training models to generalize to out-of-distribution (OOD) chemical spaces remains fundamentally constrained by the high cost of experimental validation. In antibiotic discovery, where whole-cell phenotypic high throughput screening (HTS) is resource-intensive, iterative ML-guided compound selection – or Active Learning (AL) – offers a pathway to efficiently navigate available chemical spaces. However, the algorithmic tradeoffs between prioritizing compound novelty (exploration), predicted bioactivity (exploitation), and their impact on OOD generalizability remain unresolved for noisy, whole-cell biological systems. In this work, we systematically evaluate three AL strategies for whole-cell bacterial bioactivity and benchmark their effects on model accuracy, hit rate, and OOD performance. Using retrospective simulations on Mycobacterium tuberculosis HTS data, we identify an optimal AL strategy that balances predicted hit/non-hit novelty with overall hit rate. We then integrate the strategy in a closed-loop Borrelia burgdorferi antibiotic discovery HTS campaign. The AL-guided approach successfully increased the experimental screening hit rate five-fold (from a 0.2% rate within investigator-selected plates to 1.0%). Further, when the trained model was applied in prospective in silico selection of highly diverse compounds across multiple bacterial species, the AL-trained whole-cell inhibition predictor demonstrates 53-fold enrichment over investigator-directed screening (11.0% experimental validation of predicted hits). Of these, 100% demonstrated the intended narrow spectrum activity for Borrelia burgdorferi. These results demonstrate that calibrated AL strategies can overcome data acquisition bottlenecks and train generalizable property predictors able to extrapolate to OOD molecules.
This work systematically evaluates the performance of a popular ML-based docking method, DiffDock-Pocket, on high-throughput screening (HTS) data sets derived from the PubChem BioAssay database, a premier source of bioactivity data.
Furyal Ahmed, M. Soellner, Charles L. Brooks· Journal of Chemical Informat...· 0 citations
The results show that the use of suitable encoder-regressor pairs together with embedding-level mix-up augmentation improves model generalizability without requiring SMILES-level augmentation, and could be applied more broadly to IC50 prediction for other kinase inhibitors.
Ju Hyung Lee, S. Choi, Utku Ozbulak et al.· Journal of Cheminformatics· 0 citations
How recent advances in machine learning are reshaping AMP research is examined, driving a shift from large-scale discovery toward precision-guided prediction and design and emphasizing integrated generative-predictive pipelines, interpretable models, and closed-loop experimental validation as key enablers for the development of potent, selective, and clinically viable antimicrobial therapeutics.
This review highlights the synergy between AI and HTS, emphasizing DL techniques such as convolutional neural networks for bioactivity prediction, recurrent neural networks for de novo design, and reinforcement learning for property optimization.
K. Herbetko, Katarzyna Herbetko, Magdalena Mikołajek et al.· Future Medicinal Chemistry· 0 citations
This study highlights the utility of large-scale MTL for pharmacokinetics profiling and contributes practical tools and data sets for the community, and reports a unified ChemProp-based MTL model capable of handling hundreds of continuous tasks simultaneously, which has practical advantages for model deployment and maintenance.
P. Llompart, C. Minoletti, G. Marcou et al.· Journal of Medicinal Chemist...· 0 citations
A machine-learning-assisted enzyme-engineering (MLEE) workflow that adds substrate-specific functional information to htFuncLib through an initial screening and sequencing round that may bypass the need for transition-state models and reduce the effort required for obtaining high-activity variants.
Li Wan, Mahdi Bagherpoor Helabad, Lena Fraedrich et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.