L-3,4-dihydroxyphenylalanine (L-Dopa) is an important pharmaceutical for the treatment of Parkinson’s disease and a precursor to numerous catechol-containing compounds. The flavin-dependent monooxygenase HpaBC is a promising biocatalyst for microbial L-Dopa production but exhibits limited native activity toward L-tyrosine. Although structure-based machine learning (ML) models have become increasingly popular for protein engineering, relatively few studies have systematically compared their performance or evaluated their integration into iterative engineering workflows. Here, we benchmarked multiple ML models for their ability to predict activity enhancing mutations in HpaBC. Experimentally validated single mutants were used to seed combinatorial design with EVOLVEpro, generating progressively improved higher-order variants. We next evaluated how expanding the EVOLVEpro training set with directed evolution derived variants influenced combinatorial predictions and finally explored an expanded sequence space by allowing combinations of both machine learning derived and directed evolution derived mutations. This workflow produced HpaBC variants with substantially improved activity. Although incorporating directed evolution data substantially altered EVOLVEpro’s predicted mutational trajectories, both training strategies converged on variants with comparable activities, demonstrating that distinct regions of sequence space can yield similarly optimized enzymes. Together, these results provide a systematic comparison of zero-shot ML models and establish an iterative framework for integrating machine learning with directed evolution to accelerate enzyme engineering.
Daniel Gutierrez, Isa Madrigal Harrison, Aaron L. Feller et al.· bioRxiv· 0 citations
Understanding how molecular structure encodes biological function remains a grand challenge in drug discovery. Here, we present PubCheF-1, a deep learning model that predicts literature-derived biological function directly from chemical structure. PubCheF-1 was trained on a dataset linking molecules to labels derived from the scientific articles in which they appear, a strategy that connects disparate compounds through the language used to describe their functionalities. When tasked with identifying inhibitors of β-lactamases, including enzymes considered largely refractory to inhibition, PubCheF-1 predicted structurally distinct compounds that collectively have activity against all β-lactamase classes. Furthermore, hit compounds directly bind the enzyme active site, restore antibiotic efficacy in multidrug-resistant high-priority pathogens, and demonstrate potent activity in animal infection models. Together, these findings establish that machine learning-based prediction of biological function derived from the language of scientific literature allows the identification of bioactive molecules at high hit rates, thereby accelerating therapeutic discovery.
Clayton W. Kosonocky, Nikol Kadeřábková, Kangsan Kim et al.· bioRxiv· 0 citations