A transfer learning-based predictor for acidophilic and alkalophilic proteins, trained on a curated non-redundant dataset, and an integrated pipeline for mining and engineering alkalophilic and thermophilic enzymes, combining sequence-based prediction, generative modeling, and multi-parameter virtual screening are developed.
Abstract
Extremozymes offer substantial potential as biocatalysts in industrial biotechnology, yet their identification and optimization remain challenging. Here, we developed AAEPre, a transfer learning-based predictor for acidophilic and alkalophilic proteins, trained on a curated non-redundant dataset. AAEPre achieved an average accuracy of 0.80 and outperformed conventional machine learning approaches. Based on this model, we developed an integrated pipeline for mining and engineering alkalophilic and thermophilic enzymes, combining sequence-based prediction, generative modeling, and multi-parameter virtual screening. This strategy enabled the discovery of a novel xylanase, 8E20, with optimal activity at 55 °C and pH 8.0, followed by large-scale in silico diversification to generate 1000,000 variants. Systematic screening identified the superior variant 8E20-178, which exhibits a 1.9-fold increase in catalytic activity, a shift in optimal pH from 8.0 to 10.0, and improved alkaline stability. Structural analysis suggests that strengthened hydrophobic interactions and charge redistribution contribute to its improved alkali tolerance. Notably, 8E20-178 has strong potential for practical use, including pulp biobleaching and beating. The AAEPre model now is available at http://106.8.105.46:10152/, and is free for users. Collectively, our work presents a generalizable and experimentally validated computational framework for enzyme discovery and optimization under extreme conditions.
Mycosporine-like amino acids (MAAs) are functional secondary metabolites renowned for their exceptional UV protection, antioxidant properties, and environmental resilience. In the MAA biosynthetic pathway, MysD is the pivotal enzyme mediating the chemical transition of the cyclohexenone core into a cyclohexenimine-type scaffold. This MysD-catalyzed secondary amino acid modification not only dictates the chemical diversity of MAAs but also facilitates a crucial bathochromic shift, moving the UV absorption maximum from the UVB range into the high-penetration UVA region. Despite its significance, MysD remains the rate-limiting step in the biosynthesis of iminomycosporine-like amino acids. To date, only eight MysD enzymes have been heterologously validated, with even fewer subjected to biochemical characterization, which severely restricts the use of conventional supervised machine learning for enzyme discovery. In this study, we developed an integrated, data-driven screening framework combining Sequence Similarity Networks (SSN), deep representation learning (UniRep), and Positive-Unlabeled Bagging (PU Bagging) to explore the MysD functional landscape. This pipeline effectively compressed the search space from approximately 951 unannotated homologues to a prioritized 42 candidates. Experimental validation led to the discovery of AdMysD from Aphanothece hegewaldii, which exhibited a 3-fold increase in catalytic efficiency for porphyra-334 production relative to the previously established benchmark, NlMysD. Notably, while demonstrating a primary preference for l-Thr, AdMysD displayed significant substrate promiscuity by accepting l-Ser, l-Ala, and l-Cys to produce iminomycosporine derivatives. Our findings provide a biocatalytic tool for the efficient production of MAAs and demonstrate the potential of a PU-learning-based prioritization strategy for identifying rare enzyme families with sparse functional annotations.
Protein thermostability is a critical property for both industrial and biomedical enzyme applications, yet experimental evaluation of mutation-induced stability changes remains laborious and costly. Here, we present ThermoFusion, a hybrid deep learning framework that integrates 3D protein structure embeddings from ThermoMPNN with sequence-based embeddings from the pretrained protein language model ESM2 to predict the effects of single-point mutations on protein stability (ΔΔG). ThermoFusion exhibits robust generalization, maintaining high predictive accuracy across out of distribution sequences with low identity to the training set – a scenario where many other machine learning models, including ThermoMPNN and state-of-the-art tools, perform poorly due to reliance on memorization. Benchmarking on a curated enzyme dataset comprising of 105 enzymes and 3144 mutations shows that ThermoFusion reliably identifies stabilizing mutations while accurately predicting stability for enzymes beyond its training set. These results establish ThermoFusion as a powerful tool for rational enzyme design beyond its training set.
Yao Wei, I. Eberini, Fabian Meyer· bioRxiv· 0 citations
gem-Difluorophosphonates are pivotal structural motifs in pharmaceuticals and bioactive molecules. While photoenzymatic catalysis provides a powerful platform to overcome the challenges of enantioselective synthesis, engineering enzymes for non-natural transformations remains an arduous, labor-intensive process. Although predictive methods utilizing protein language models (PLMs) offer fitness landscape guidance, they often struggle to generalize across diverse protein families or accurately map sequence to catalytic activity. Here, we report a small-sample, accelerated evolution strategy that integrates focused rational iterative site-specific mutagenesis (FRISM) with the EVOLVEpro model. This synergistic approach identifies high-activity and enantiospecific variants through structure-based hotspot identification and active learning, requiring minimal experimental throughput. By screening only 40 variants over three evolutionary rounds, we identified four beneficial mutations whose combinations enable the synthesis of diverse fluorinated products with up to > 99% yield and 98:2 enantiomeric ratio (e.r.)-a 65% reduction in workload compared to exhaustive screening. Mechanistic investigations suggest an electron donor-acceptor (EDA)-complex-free radical addition pathway, terminated by the flavin semiquinone (FMNsq) or the active-site residue Y343. This study provides a robust, "lightweight" machine learning framework for the rapid development of new-to-nature photoenzymatic transformations.
Reliable estimation of downstream performance in low-data peptide machine learning is critical for guiding early-stage AI-driven peptide engineering. Yet, it is often unclear how to assess whether a model will be effective in iterative discovery settings. Here, we show that the cross validation R² score can serve as a simple and robust proxy for predicting active learning workflow performance, enabling early-stage evaluation of model suitability for sequential peptide optimization. To support this, we introduce SCARSE, a machine learning framework combining ESM-2 protein language model embeddings with Gaussian process regression and extremely randomized trees classification, designed for low-resource peptide property prediction (20–500 training samples). We benchmark SCARSE across 23 peptide and small-protein datasets covering substitution and indel variants, antimicrobial peptides, cell-penetrating peptides, and toxic/non-toxic peptides. SCARSE significantly outperforms a hand-engineered descriptor baseline on substitution and indel tasks, while comparable performance was achieved on shorter peptide non-mutant datasets where simpler descriptors capture enough of the signal. In simulated active learning workflows, SCARSE consistently outperforms baseline and random sampling strategies. Notably, we demonstrate that CV R² computed from as few as 50 labeled peptides can be sufficient to estimate final active learning end-point performance, providing a practical, data-efficient criterion for deciding whether a given dataset combined with SCARSE is suitable for iterative peptide discovery. SCARSE is released as a pip package and is available via HuggingFace Spaces to facilitate integration into peptide engineering workflows.
Leo Andrekson, Robin Rydbergh, Rocío Mercado et al.· bioRxiv· 0 citations
Nowadays, enzyme engineering has moved from traditional structure-based mutagenesis and directed evolution to data-intensive, AI-assisted design paradigms that involve the rapid discovery and optimization of biocatalysts. Whereas classical approaches relied on rational design and experimental screening, advances in high-throughput sequencing, modeling, and machine learning have enabled predictive exploration of sequence-structure-function relationships in enzymes. Importantly, the latest protein language models and deep learning approaches enable accurate prediction of mutational outcomes, stability engineering, and functional annotation at an unprecedented scale. Generative AI models also enable the design of novel enzymes by predicting protein sequences with tailored catalytic functions and broadened substrate specificity. AI combined with design-build-test-learn (DBTL) automation and synthetic biology has enabled the creation of closed-loop engineering workflows for rapid, iterative optimization. This review examines enzyme engineering from classical methods to AI-assisted biocatalyst development, highlighting key advances, challenges, and emerging trends in autonomous laboratories, sustainable biocatalysis, and computational protein design.
Mati Ullah, Muhammad Rizwan, Vivian Andoh et al.· Journal of Agricultural and...· 0 citations
The rapid emergence of metallo-b-lactamase-mediated antibiotic resistance has created an urgent need for new inhibitor discovery strategies. In this work, a machine-learning-guided workflow was developed to generate and prioritize potential inhibitors targeting NDM-1. A SMILES-based variational autoencoder was first pretrained on a broad molecular dataset to learn general chemical syntax and latent molecular representations. The model was then fine-tuned on an 8-hydroxyquinoline-enriched dataset to bias molecular generation toward zinc-binding chemical space relevant to metallo-β-lactamase inhibition. Generated compounds were processed through structural filtering and docking-based evaluation to create training data for downstream predictive modeling. Molecular fingerprints and physicochemical descriptors were then used to train XGBoost models for docking score prediction and classification of potential binders. Classification proved especially useful for prescreening because it avoided overinterpreting small differences in noisy docking scores while still enriching for compounds likely to perform well in docking. The resulting workflow demonstrates how generative modeling and supervised machine learning can be combined to reduce chemical search space, prioritize candidate inhibitors, and guide computational drug discovery. Although experimental validation remains necessary, this approach provides a scalable framework for identifying promising zinc-binding compounds for further molecular simulation and inhibitor development that can be expanded in future studies.
Anthony M. Baudino, Kari L. Stone· AI Chemistry· 0 citations