Back to feed
Open access

Machine Learning-Assisted Evolution of Broadly Functional Enzyme Libraries

Jul 2026 · bioRxiv · 0 citations · 56 references
Biology

TL;DR

Results indicate that supervised machine learning can help guide the construction of high-value enzyme libraries with expanded catalytic scope, and suggest that supervised machine learning can help guide the construction of high-value enzyme libraries with expanded catalytic scope.

Abstract

Biocatalysis offers sustainable solutions to pressing challenges in chemical synthesis by exploiting the remarkable efficiency and selectivity of enzymes. Importantly, enzymes are able to accommodate non-native substrates and mediate transformations outside of their natural repertoire. Enzymes can be engineered for diverse applications by harnessing these ‘promiscuous’ activities and optimizing them using directed evolution (DE). The success of a DE campaign, however, depends on the availability of a protein starting point that displays detectable levels of the desired function. To find a starting point, researchers often screen libraries of protein variants for novel activities, typically with low rates of success. Here, instead, we diversified the active site of a desirable ‘parent’ protein and applied machine learning to generate informed, promiscuous libraries of protein variants. Specifically, we tested 26 different carbene and nitrene transfer reactions and used active learning-assisted directed evolution (ALDE) to generate optimized protoglobin variants with high activity across multiple reactions. We observed improvements in activity and selectivity for every reaction performed by the parent enzyme in at least one member of the ALDE-predicted libraries. Moreover, variants from these libraries can catalyze 5 out of 10 reactions not catalyzed by the parent protoglobin. These results indicate that supervised machine learning can help guide the construction of high-value enzyme libraries with expanded catalytic scope.

Read PDF

Similar papers

Open access Jul 2026

Combining Stability-Centered Atomistic Design with Machine Learning for Targeted Enzyme Optimization

FuncLib and high-throughput FuncLib (htFuncLib) generate diverse, functional protein libraries using a stability-centered design; however, this substrate-independent approach lacks target-specific functional constraints. We developed a machine-learning-assisted enzyme-engineering (MLEE) workflow that adds substrate-specific functional information to htFuncLib through an initial screening and sequencing round. The system was benchmarked using previously published four-position fitness landscapes of three different proteins. The MLEE workflow successfully generated compact libraries enriched in globally high-fitness variants. After the initial training phase, an MLEE-enriched library of just 12 variants increased the hit rate for the global top-0.05% variants by 5- to 12-fold relative to the htFuncLib baseline. Screening a larger set of 96 variants recovered at least one of these top-performing enzymes in 61.3–99.4% of the simulations. We then applied MLEE to MthUPO-catalyzed β-damascone hydroxylation. Across two rounds, 506 distinct variants were screened and sequenced. While the initial substrate-independent htFuncLib library yielded 14% of variants with activity above the wild type, the MLEE-enriched library increased this hit rate to 90% (97 of 108 variants) with activity above the wild type. The best variant increased the turnover number for 4-hydroxy-β-damascone by 11.8-fold and achieved >99% regioisomeric excess. MLEE may bypass the need for transition-state models and reduce the effort required for obtaining high-activity variants. TABLE OF CONTENT

Li Wan, Mahdi Bagherpoor Helabad, Lena Fraedrich et al. · 0 citations
Review Aug 2026

Harnessing Machine Learning for Enzyme Enantioselectivity: Toward Intelligent Biocatalysis Engineering

Enzyme enantioselectivity remains a central challenge in the biocatalytic synthesis of chiral pharmaceuticals and fine chemicals. Conventional evolution methods, constrained by prolonged experimental cycles and computationally intensive processes, struggle to comprehensively explore complex protein sequences. Although machine learning (ML) has been successfully applied to optimize enzymatic properties, such as catalytic efficiency and stability, its potential in enantioselectivity engineering remains underexploited. This review systematically evaluates how machine learning integrates multimodal datasets, including sequence, structural, and functional performance data, to advance the discovery and engineering of naturally enantioselective enzymes, and enable the de novo design of artificial enzymes for reaction-relevant applications. Furthermore, key algorithmic frameworks are reviewed. Persistent challenges such as data heterogeneity and limited model generalization capabilities are also critically examined. Moreover, machine learning is increasingly expected to bridge molecular-level enantioselective design with pathway optimization and reactor-scale process engineering, enabling a more integrated and efficient chiral biomanufacturing pipeline. Overall, these advancements open new avenues for efficient chiral biosynthesis and provide theoretical foundations and technical blueprints for a paradigm shift toward intelligence-driven biocatalysis engineering.

Jie Gu, Yan Xu, Xiaoyan Sun et al. · 0 citations
Review Open access Jul 2026

Applications and limitations of AI tools in enzyme design

Enzymes catalyze various different chemical reactions often with high efficiency and selectivity compared to synthetic catalysts. Advances in protein engineering over the past decades have allowed researchers to design enzymes, improving their catalytic performance and adapting them for specific or entirely novel chemical reactions. However, the need for the experimental validation of thousands of computational designs remains one of the major bottlenecks. Yet another restriction is our still limited knowledge about transition state architectures, effects of mutations, active‐site dynamics just to name a few. To overcome this, the combination of experimental and computational methods is essential, yet many experimentalists are facing significant obstacles when entering the field of computational enzyme design. To address these obstacles, this review offers a comprehensive introduction and overview of several current artificial intelligence (AI)‐driven methods available for enzyme design, with a focus on reaction‐to‐sequence design, structure prediction, substrate scope prediction, engineering of stable variants, design of enzymes with non‐canonical amino acids, and de novo design. Subsequently, this work serves as an accessible guide for experimental researchers with interest in learning how to use AI‐based computational methods in enzyme engineering.

Rosa Teijeiro-Juiz, Nina Egeler, Grzegorz Jamróg et al. · 1 citation
Review Open access Jun 2026

Machine Learning-Guided Enzyme Engineering Approaches for Enhanced Biocatalytic Efficiency: Concepts, Mechanisms, and Future Directions

Biocatalysis has emerged as a mainstay in the field of sustainable chemical synthesis owing to its high selectivity, mild reaction conditions, and reduced environmental impact. Traditional enzyme engineering approaches, such as rational design and directed evolution, are often associated with limited throughput and a limited understanding of sequence–structure–function relationships, despite high experimental costs. In recent years, the integration of machine learning (ML) into enzyme engineering has emerged as a transformative approach, enabling data-driven prediction, design, and optimization of biocatalysts, thereby enhancing performance and applications. This review provides a comprehensive overview of ML-guided strategies to improve key enzymatic parameters, including the turnover number (kcat), substrate affinity (Km), and catalytic efficiency (kcat/Km), with a focus on mechanistic insights and performance outcomes. The integration of ML models into design–build–test–learn (DBTL) cycles accelerated directed evolution, reduced screening efforts, and enabled targeted mutagenesis. Beyond applications, this review also discusses the current limitations of ML-guided approaches, including data scarcity, model interpretability, and challenges in predicting complex mutations and allosteric effects. The gap between computational predictions and experimental outcomes is identified, and the role of ML integration with enzyme kinetics, molecular dynamics, and high-throughput experimentation is emphasized. Future directions, such as generative AI, explainable ML, and autonomous laboratories, are discussed for next-generation biocatalytic applications.

W. Ahsan · 0 citations
Review Aug 2026

Enzyme Engineering: From Classical Strategies to AI-Driven Biocatalyst Design.

Nowadays, enzyme engineering has moved from traditional structure-based mutagenesis and directed evolution to data-intensive, AI-assisted design paradigms that involve the rapid discovery and optimization of biocatalysts. Whereas classical approaches relied on rational design and experimental screening, advances in high-throughput sequencing, modeling, and machine learning have enabled predictive exploration of sequence-structure-function relationships in enzymes. Importantly, the latest protein language models and deep learning approaches enable accurate prediction of mutational outcomes, stability engineering, and functional annotation at an unprecedented scale. Generative AI models also enable the design of novel enzymes by predicting protein sequences with tailored catalytic functions and broadened substrate specificity. AI combined with design-build-test-learn (DBTL) automation and synthetic biology has enabled the creation of closed-loop engineering workflows for rapid, iterative optimization. This review examines enzyme engineering from classical methods to AI-assisted biocatalyst development, highlighting key advances, challenges, and emerging trends in autonomous laboratories, sustainable biocatalysis, and computational protein design.

Mati Ullah, Muhammad Rizwan, Vivian Andoh et al. · 0 citations