Back to feed

Machine learning-assisted discovery of AdMysD for enhanced porphyra-334 biosynthesis.

Jul 2026 · Bioorganic chemistry (Print) · Vol 181, pp. 110312 · 0 citations · 43 references
Medicine

Abstract

Mycosporine-like amino acids (MAAs) are functional secondary metabolites renowned for their exceptional UV protection, antioxidant properties, and environmental resilience. In the MAA biosynthetic pathway, MysD is the pivotal enzyme mediating the chemical transition of the cyclohexenone core into a cyclohexenimine-type scaffold. This MysD-catalyzed secondary amino acid modification not only dictates the chemical diversity of MAAs but also facilitates a crucial bathochromic shift, moving the UV absorption maximum from the UVB range into the high-penetration UVA region. Despite its significance, MysD remains the rate-limiting step in the biosynthesis of iminomycosporine-like amino acids. To date, only eight MysD enzymes have been heterologously validated, with even fewer subjected to biochemical characterization, which severely restricts the use of conventional supervised machine learning for enzyme discovery. In this study, we developed an integrated, data-driven screening framework combining Sequence Similarity Networks (SSN), deep representation learning (UniRep), and Positive-Unlabeled Bagging (PU Bagging) to explore the MysD functional landscape. This pipeline effectively compressed the search space from approximately 951 unannotated homologues to a prioritized 42 candidates. Experimental validation led to the discovery of AdMysD from Aphanothece hegewaldii, which exhibited a 3-fold increase in catalytic efficiency for porphyra-334 production relative to the previously established benchmark, NlMysD. Notably, while demonstrating a primary preference for l-Thr, AdMysD displayed significant substrate promiscuity by accepting l-Ser, l-Ala, and l-Cys to produce iminomycosporine derivatives. Our findings provide a biocatalytic tool for the efficient production of MAAs and demonstrate the potential of a PU-learning-based prioritization strategy for identifying rare enzyme families with sparse functional annotations.

View source

Similar papers

Jul 2026

Machine-Learning-Enabled Rapid Evolution of Photoenzymes for the Asymmetric Synthesis of gem-Difluorophosphonates.

gem-Difluorophosphonates are pivotal structural motifs in pharmaceuticals and bioactive molecules. While photoenzymatic catalysis provides a powerful platform to overcome the challenges of enantioselective synthesis, engineering enzymes for non-natural transformations remains an arduous, labor-intensive process. Although predictive methods utilizing protein language models (PLMs) offer fitness landscape guidance, they often struggle to generalize across diverse protein families or accurately map sequence to catalytic activity. Here, we report a small-sample, accelerated evolution strategy that integrates focused rational iterative site-specific mutagenesis (FRISM) with the EVOLVEpro model. This synergistic approach identifies high-activity and enantiospecific variants through structure-based hotspot identification and active learning, requiring minimal experimental throughput. By screening only 40 variants over three evolutionary rounds, we identified four beneficial mutations whose combinations enable the synthesis of diverse fluorinated products with up to > 99% yield and 98:2 enantiomeric ratio (e.r.)-a 65% reduction in workload compared to exhaustive screening. Mechanistic investigations suggest an electron donor-acceptor (EDA)-complex-free radical addition pathway, terminated by the flavin semiquinone (FMNsq) or the active-site residue Y343. This study provides a robust, "lightweight" machine learning framework for the rapid development of new-to-nature photoenzymatic transformations.

Hongkui Wang, Jiafan Xu, Jiahai Zhou et al. · 0 citations
Open access Jul 2026

A Machine Learning-Based Genome Mining Approach Reveals Unprecedented Biarylitide Diversity

Biarylitides are a group of bacterial ribosomally synthesized and post-translationally modified peptides (RiPPs) that contain a biaryl bridge formed by dedicated cytochrome P450 enzymes that can introduce different cross-links. The biarylitides are produced via a five-amino-acid precursor peptide, encoded by a minimal 18 bp gene that evades automatic detection. Previous genome mining approaches for biarylitides do not capture their full biosynthetic space. We therefore repurposed a machine learning algorithm to comprehensively chart the biosynthetic space of the biarylitides, including variation of precursor motifs, P450, and additional modifying enzymes, which yielded 277 biarylitide biosynthetic gene clusters (BGCs). We experimentally investigated biaryl formation with previously uninvestigated core peptide motifs, including YWH, YVH, and YWY, and elucidated the nature of these cross-links. This study significantly expands the biarylitide precursor and BGC diversity and provides directions for the systematic exploration of other RiPP families.

Leo Padva, Jemma Gullick, Friederike Biermann et al. · 0 citations
Jul 2026

A deep learning and generative modeling pipeline for mining and engineering alkaline-stsable xylanases.

Extremozymes offer substantial potential as biocatalysts in industrial biotechnology, yet their identification and optimization remain challenging. Here, we developed AAEPre, a transfer learning-based predictor for acidophilic and alkalophilic proteins, trained on a curated non-redundant dataset. AAEPre achieved an average accuracy of 0.80 and outperformed conventional machine learning approaches. Based on this model, we developed an integrated pipeline for mining and engineering alkalophilic and thermophilic enzymes, combining sequence-based prediction, generative modeling, and multi-parameter virtual screening. This strategy enabled the discovery of a novel xylanase, 8E20, with optimal activity at 55 °C and pH 8.0, followed by large-scale in silico diversification to generate 1000,000 variants. Systematic screening identified the superior variant 8E20-178, which exhibits a 1.9-fold increase in catalytic activity, a shift in optimal pH from 8.0 to 10.0, and improved alkaline stability. Structural analysis suggests that strengthened hydrophobic interactions and charge redistribution contribute to its improved alkali tolerance. Notably, 8E20-178 has strong potential for practical use, including pulp biobleaching and beating. The AAEPre model now is available at http://106.8.105.46:10152/, and is free for users. Collectively, our work presents a generalizable and experimentally validated computational framework for enzyme discovery and optimization under extreme conditions.

Ruohan Zhang, Yiyang Zhang, Zhonghao Deng et al. · 0 citations
Open access Jul 2026

Machine Learning Integrated Designing and Screening of 8-Hydroxyquinoline Based Metallo-β-Lactamase Inhibitors

The rapid emergence of metallo-b-lactamase-mediated antibiotic resistance has created an urgent need for new inhibitor discovery strategies. In this work, a machine-learning-guided workflow was developed to generate and prioritize potential inhibitors targeting NDM-1. A SMILES-based variational autoencoder was first pretrained on a broad molecular dataset to learn general chemical syntax and latent molecular representations. The model was then fine-tuned on an 8-hydroxyquinoline-enriched dataset to bias molecular generation toward zinc-binding chemical space relevant to metallo-β-lactamase inhibition. Generated compounds were processed through structural filtering and docking-based evaluation to create training data for downstream predictive modeling. Molecular fingerprints and physicochemical descriptors were then used to train XGBoost models for docking score prediction and classification of potential binders. Classification proved especially useful for prescreening because it avoided overinterpreting small differences in noisy docking scores while still enriching for compounds likely to perform well in docking. The resulting workflow demonstrates how generative modeling and supervised machine learning can be combined to reduce chemical search space, prioritize candidate inhibitors, and guide computational drug discovery. Although experimental validation remains necessary, this approach provides a scalable framework for identifying promising zinc-binding compounds for further molecular simulation and inhibitor development that can be expanded in future studies.

Anthony M. Baudino, Kari L. Stone · 0 citations
Open access Jul 2026

Machine learning guided cell-free expression maps the biochemical landscape of carbonic anhydrase

This work demonstrates that integrating cell-free enzyme engineering with machine learning enables opportunities for high-throughput experimental measurements to benchmark and improve protein language models, accelerate design loops, and expand functional exploration within protein families where experimental information is limited.

J. Lazar, Evan Komp, I. Martínez et al. · 1 citation
Open access Aug 2026

Systematic prioritization of candidate genes in camptothecin biosynthesis using multi-omics and deep learning.

Camptothecin (CPT), a plant-derived monoterpene indole alkaloid first identified in Camptotheca acuminata, is a drug precursor widely used for cancer chemotherapeutics. However, the full set of genes responsible for CPT biosynthesis remains unclear, hindering efforts to elucidate the complete pathway or establish biosynthetic production of CPT in heterologous hosts. In this study, we engineered an experimental callus system for inducible production of CPT, which enabled multi-omics and deep learning analyses to identify candidate genes in CPT biosynthesis. We first generated an improved genome assembly and gene annotation for C. acuminata. We then leveraged the natural variation of CPT levels in C. acuminata tissues and performed transcriptomic analysis of multiple callus and tissue types to shortlist candidate enzymes responsible for CPT biosynthesis. Finally, we conducted large-scale deep learning-enabled protein-ligand complex structure prediction to prioritize 117 candidate enzymes for studies that map their roles in CPT biochemical reactions. By integrating experimental, genomic, transcriptomic, and deep learning approaches, this study provides a valuable foundation for the complete elucidation of the CPT biosynthetic pathway.

Shenqiu Wang, Xing Wu, Maria Moreno et al. · 0 citations