Back to feed
Open access

Machine learning guided cell-free expression maps the biochemical landscape of carbonic anhydrase

Jul 2026 · bioRxiv · 1 citation · 12 references
Biology

TL;DR

This work demonstrates that integrating cell-free enzyme engineering with machine learning enables opportunities for high-throughput experimental measurements to benchmark and improve protein language models, accelerate design loops, and expand functional exploration within protein families where experimental information is limited.

Abstract

Carbonic anhydrases are among the fastest known biocatalysts, reversibly facilitating the hydration of CO2 to HCO3- at rates up to 107 s-1, which warrants their investigation for industrial carbon capture technologies. However, engineering carbonic anhydrases to maintain stability under harsh industrial process conditions remains a key challenge, and sequence-to-function datasets compatible with machine learning to inform forward engineering are lacking. Here, we developed a high-throughput platform that couples cell-free gene expression with a gaseous CO2 colorimetric assay to map the fitness landscapes of carbonic anhydrases. From 96 diverse natural homologs, we identified a robust variant from the Aquificota phylum and conducted an exhaustive mutational scan and functional assessment of this enzyme at 70°C and 90°C, covering >99% of all single-amino acid substitutions (totaling 4,365 mutations assayed in 39,285 reactions). This biochemical landscape was used to benchmark 22 zero-shot protein fitness models and identify critical mutations that improved enzyme stability at 90°C by more than three-fold. We then used both zero-shot protein language models and supervised learning to filter 419 model-generated variants from a ProteinMPNN library of 100,000 sequences, leading to a best-in-class enzyme that retained activity after incubation at 95°C. This work demonstrates that integrating cell-free enzyme engineering with machine learning enables opportunities for high-throughput experimental measurements to benchmark and improve protein language models, accelerate design loops, and expand functional exploration within protein families where experimental information is limited.

Read PDF

Similar papers

Open access Jul 2026

Machine learning-assisted directed evolution of plant Rubisco

Ribulose-1,5-bisphosphate carboxylase/oxygenase (Rubisco) is foundational to life on Earth, catalyzing carbon dioxide (CO2) fixation to generate biomass. However, Rubisco is a slow and inefficient enzyme that has proven challenging to engineer. We applied the structure-informed machine learning (ML) model ESM-IF1 to identify plausible amino acid sites in the large subunit of Nicotiana tabacum Rubisco to target for directed evolution. ML-assisted library design followed by selection in Rubisco-dependent Escherichia coli identified multiple enriched variants displaying improved catalytic efficiency. Several improved variants carried amino acid changes not found in the evolutionary lineage of plants, despite being assembly competent in plant chloroplasts, demonstrating that ML-assisted protein design can explore functional sequence space beyond what is observed from natural sequence diversity. Most prominently, the T391I substitution improved carboxylation rate by 29% and aerobic carboxylation efficiency by 43%. Our findings demonstrate the utility of ML-assisted evolution for engineering Rubisco with improved carboxylation efficiency and potential for enhancing crop productivity.

Julie L. McDonald, Jiachen Lin, Yunlong Zhao et al. · 0 citations
Open access Jul 2026

Combining Stability-Centered Atomistic Design with Machine Learning for Targeted Enzyme Optimization

FuncLib and high-throughput FuncLib (htFuncLib) generate diverse, functional protein libraries using a stability-centered design; however, this substrate-independent approach lacks target-specific functional constraints. We developed a machine-learning-assisted enzyme-engineering (MLEE) workflow that adds substrate-specific functional information to htFuncLib through an initial screening and sequencing round. The system was benchmarked using previously published four-position fitness landscapes of three different proteins. The MLEE workflow successfully generated compact libraries enriched in globally high-fitness variants. After the initial training phase, an MLEE-enriched library of just 12 variants increased the hit rate for the global top-0.05% variants by 5- to 12-fold relative to the htFuncLib baseline. Screening a larger set of 96 variants recovered at least one of these top-performing enzymes in 61.3–99.4% of the simulations. We then applied MLEE to MthUPO-catalyzed β-damascone hydroxylation. Across two rounds, 506 distinct variants were screened and sequenced. While the initial substrate-independent htFuncLib library yielded 14% of variants with activity above the wild type, the MLEE-enriched library increased this hit rate to 90% (97 of 108 variants) with activity above the wild type. The best variant increased the turnover number for 4-hydroxy-β-damascone by 11.8-fold and achieved >99% regioisomeric excess. MLEE may bypass the need for transition-state models and reduce the effort required for obtaining high-activity variants. TABLE OF CONTENT

Li Wan, Mahdi Bagherpoor Helabad, Lena Fraedrich et al. · 0 citations
Open access Jul 2026

Machine Learning-Assisted Evolution of Broadly Functional Enzyme Libraries

Results indicate that supervised machine learning can help guide the construction of high-value enzyme libraries with expanded catalytic scope, and suggest that supervised machine learning can help guide the construction of high-value enzyme libraries with expanded catalytic scope.

Ravi G. Lal, Jason Yang, Ziyan Zhang et al. · 0 citations
Jul 2026

Machine-Learning-Enabled Rapid Evolution of Photoenzymes for the Asymmetric Synthesis of gem-Difluorophosphonates.

gem-Difluorophosphonates are pivotal structural motifs in pharmaceuticals and bioactive molecules. While photoenzymatic catalysis provides a powerful platform to overcome the challenges of enantioselective synthesis, engineering enzymes for non-natural transformations remains an arduous, labor-intensive process. Although predictive methods utilizing protein language models (PLMs) offer fitness landscape guidance, they often struggle to generalize across diverse protein families or accurately map sequence to catalytic activity. Here, we report a small-sample, accelerated evolution strategy that integrates focused rational iterative site-specific mutagenesis (FRISM) with the EVOLVEpro model. This synergistic approach identifies high-activity and enantiospecific variants through structure-based hotspot identification and active learning, requiring minimal experimental throughput. By screening only 40 variants over three evolutionary rounds, we identified four beneficial mutations whose combinations enable the synthesis of diverse fluorinated products with up to > 99% yield and 98:2 enantiomeric ratio (e.r.)-a 65% reduction in workload compared to exhaustive screening. Mechanistic investigations suggest an electron donor-acceptor (EDA)-complex-free radical addition pathway, terminated by the flavin semiquinone (FMNsq) or the active-site residue Y343. This study provides a robust, "lightweight" machine learning framework for the rapid development of new-to-nature photoenzymatic transformations.

Hongkui Wang, Jiafan Xu, Jiahai Zhou et al. · 0 citations
Jul 2026

Machine learning-assisted discovery of AdMysD for enhanced porphyra-334 biosynthesis.

Mycosporine-like amino acids (MAAs) are functional secondary metabolites renowned for their exceptional UV protection, antioxidant properties, and environmental resilience. In the MAA biosynthetic pathway, MysD is the pivotal enzyme mediating the chemical transition of the cyclohexenone core into a cyclohexenimine-type scaffold. This MysD-catalyzed secondary amino acid modification not only dictates the chemical diversity of MAAs but also facilitates a crucial bathochromic shift, moving the UV absorption maximum from the UVB range into the high-penetration UVA region. Despite its significance, MysD remains the rate-limiting step in the biosynthesis of iminomycosporine-like amino acids. To date, only eight MysD enzymes have been heterologously validated, with even fewer subjected to biochemical characterization, which severely restricts the use of conventional supervised machine learning for enzyme discovery. In this study, we developed an integrated, data-driven screening framework combining Sequence Similarity Networks (SSN), deep representation learning (UniRep), and Positive-Unlabeled Bagging (PU Bagging) to explore the MysD functional landscape. This pipeline effectively compressed the search space from approximately 951 unannotated homologues to a prioritized 42 candidates. Experimental validation led to the discovery of AdMysD from Aphanothece hegewaldii, which exhibited a 3-fold increase in catalytic efficiency for porphyra-334 production relative to the previously established benchmark, NlMysD. Notably, while demonstrating a primary preference for l-Thr, AdMysD displayed significant substrate promiscuity by accepting l-Ser, l-Ala, and l-Cys to produce iminomycosporine derivatives. Our findings provide a biocatalytic tool for the efficient production of MAAs and demonstrate the potential of a PU-learning-based prioritization strategy for identifying rare enzyme families with sparse functional annotations.

Longfei Yuan, Shuting Feng, Sheng Xie et al. · 0 citations