Skip to content

Similar papers

Open access Jul 2026

Machine Learning-Assisted Evolution of Broadly Functional Enzyme Libraries

Results indicate that supervised machine learning can help guide the construction of high-value enzyme libraries with expanded catalytic scope, and suggest that supervised machine learning can help guide the construction of high-value enzyme libraries with expanded catalytic scope.

Ravi G. Lal, Jason Yang, Ziyan Zhang et al. · 0 citations
Open access Jul 2026

Combining Stability-Centered Atomistic Design with Machine Learning for Targeted Enzyme Optimization

FuncLib and high-throughput FuncLib (htFuncLib) generate diverse, functional protein libraries using a stability-centered design; however, this substrate-independent approach lacks target-specific functional constraints. We developed a machine-learning-assisted enzyme-engineering (MLEE) workflow that adds substrate-specific functional information to htFuncLib through an initial screening and sequencing round. The system was benchmarked using previously published four-position fitness landscapes of three different proteins. The MLEE workflow successfully generated compact libraries enriched in globally high-fitness variants. After the initial training phase, an MLEE-enriched library of just 12 variants increased the hit rate for the global top-0.05% variants by 5- to 12-fold relative to the htFuncLib baseline. Screening a larger set of 96 variants recovered at least one of these top-performing enzymes in 61.3–99.4% of the simulations. We then applied MLEE to MthUPO-catalyzed β-damascone hydroxylation. Across two rounds, 506 distinct variants were screened and sequenced. While the initial substrate-independent htFuncLib library yielded 14% of variants with activity above the wild type, the MLEE-enriched library increased this hit rate to 90% (97 of 108 variants) with activity above the wild type. The best variant increased the turnover number for 4-hydroxy-β-damascone by 11.8-fold and achieved >99% regioisomeric excess. MLEE may bypass the need for transition-state models and reduce the effort required for obtaining high-activity variants. TABLE OF CONTENT

Li Wan, Mahdi Bagherpoor Helabad, Lena Fraedrich et al. · 0 citations
Open access Jul 2026

An enzyme-specific protein language model for catalytic property prediction

Enzymes drive cellular metabolism, yet predicting catalytic properties from amino acid sequences remains challenging. Existing protein language models (PLMs) provide powerful general-purpose representations but are often inefficient for high-throughput screening and insufficiently adapted to enzyme-specific tasks. Here, we propose EnzGFM, an enzyme-specific PLM based on a Mamba-Transformer hybrid architecture with hierarchical pre-training to capture enzyme-specific patterns. Across enzyme property prediction benchmarks, EnzGFM consistently outperforms Transformer-based PLMs with 2–5-fold acceleration, achieving relative improvements of 16.67% in kinetic parameter prediction, 15.69% in enzyme-reaction mapping, 13.19% in EC number classification, and 20.04% in mutation effect assessment. Building on EnzGFM, we develop EnzGFM-Agent, an enzyme-focused agentic pipeline. Experimental validation further suggests that EnzGFM-Agent can enrich beneficial variants within small candidate pools. Together, these results demonstrate that EnzGFM captures enzyme-specific sequence-function patterns, while EnzGFM-Agent translates these predictions into experimentally actionable candidates and can help reduce wet-lab screening burden for practical enzyme engineering. Enzyme function prediction from amino acid sequences remains a central challenge in computational biology, despite recent advances in protein language models. This manuscript introduces EnzGFM, an enzyme-specific hybrid model that improves both accuracy and efficiency across multiple prediction tasks and, together with the EnzGFM-Agent pipeline, demonstrates the ability to identify experimentally validated beneficial variants while reducing screening effort.

Chong Wang, Mengyao Li, Shaolei Geng et al. · 0 citations
Review Aug 2026

Enzyme Engineering: From Classical Strategies to AI-Driven Biocatalyst Design.

Nowadays, enzyme engineering has moved from traditional structure-based mutagenesis and directed evolution to data-intensive, AI-assisted design paradigms that involve the rapid discovery and optimization of biocatalysts. Whereas classical approaches relied on rational design and experimental screening, advances in high-throughput sequencing, modeling, and machine learning have enabled predictive exploration of sequence-structure-function relationships in enzymes. Importantly, the latest protein language models and deep learning approaches enable accurate prediction of mutational outcomes, stability engineering, and functional annotation at an unprecedented scale. Generative AI models also enable the design of novel enzymes by predicting protein sequences with tailored catalytic functions and broadened substrate specificity. AI combined with design-build-test-learn (DBTL) automation and synthetic biology has enabled the creation of closed-loop engineering workflows for rapid, iterative optimization. This review examines enzyme engineering from classical methods to AI-assisted biocatalyst development, highlighting key advances, challenges, and emerging trends in autonomous laboratories, sustainable biocatalysis, and computational protein design.

Mati Ullah, Muhammad Rizwan, Vivian Andoh et al. · 0 citations
Open access Aug 2026

Coevolution-informed Bayesian optimization for sample-efficient protein design

This work introduces ALSEBO (Active Learning Sequence Exploration via Bayesian Optimization), which couples a generative latent sequence landscape to Bayesian optimization and featurizes candidates with direct-coupling-analysis (DCA) coevolutionary statistics.

D. P. Kulathunga, Divyanshu Shukla, D. Potoyan · 0 citations
Open access Jul 2026

AUKAT: Conditional VAE-Driven Augmentation and Neural Modeling of Enzyme Turnover Numbers

Accurate prediction of enzyme turnover numbers (kcat) is essential for applications in systems biology, metabolic engineering, and drug discovery, yet remains challenging due to the limited availability and uneven distribution of experimental data. Here, we present AUKAT, an integrated framework that combines conditional generative modeling with deep neural prediction to improve kcat estimation. A conditional variational autoencoder generates synthetic training instances in embedding space, followed by a selection pipeline that retains samples with strong agreement across independent evaluators, thereby ensuring data reliability. A hybrid convolutional neural network and transformer-based architecture is then used to predict kcat from substrate, enzyme functional, and species embeddings. Incorporating synthetic data improved predictive performance for both random forest and neural network models in five-fold cross-validation, with larger gains observed for the neural network architecture. Benchmarking against DLKcat demonstrated comparable predictive accuracy on the standard test set, while evaluation on stricter unseen subsets indicated improved generalization for low-similarity substrates and enzymes. Feature importance analysis further showed that AUKAT leverages substrate, enzyme functional, and species information in a more balanced manner rather than relying predominantly on a single feature source. In addition, AUKAT-human, a specialized model trained using a pre-training and fine-tuning strategy, achieved improved prediction accuracy for human enzyme kinetics. Overall, AUKAT provides a scalable approach for enzyme kinetics prediction and offers a practical solution to data scarcity in biochemical modeling.

Mengmeng Liu, Xialong Ni, Michal Brylinski · 0 citations