The exponential growth in available protein sequence data has broadened enzyme discovery opportunities but simultaneously highlighted a significant gap between sequence and function information. Traditional tools like phylogenetic trees and sequence similarity networks (SSNs) are widely adopted for sampling enzymes for novel transformations. However, their utility suffers from inherent limitations, which are exacerbated for large enzyme families. Phylogenetic trees, while useful for studying evolutionary relationships, become computationally intensive and difficult to visualize for larger protein datasets. SSNs, on the other hand, are sensitive to user-defined thresholds for sequence clustering and easily fail to capture more distant relationships between clusters. Additionally, both tools are alignment based and cannot capture higher-order interactions between residues. In this study, we address these limitations by optimizing a variational autoencoder (VAE)-based latent space model to visualize and explore enzyme sequence-function landscapes. By training our models on simulated datasets and real enzyme families, such as cyclases and flavin-dependent monooxygenases (FDMOs), we demonstrated that the optimized latent space effectively preserves phylogenetic relationships and enables high-resolution clustering for functionally distinct enzymes. The models further outperform traditional SSNs in capturing local and global relationships in a continuous two-dimensional space, enabling the discovery of multiple uncharacterized FDMOs for oxidative dearomatization and decarboxylative hydroxylation that illustrates their application. Our findings show that low-dimensional latent spaces can serve as valuable tools for enzyme discovery, allowing for interpolation and extrapolation to guide novel enzyme sampling for biocatalytic reactions.
Chang-Hwa Chiang, Daniel Ong, A. Narayan et al.· Proceedings of the National...· 0 citations
Achieving selective oxygenation chemistry offers a powerful strategy for streamlining routes to complex target molecules, yet the development of site-selective C-H functionalization reactions is challenging. Several enzyme families are capable of catalyzing C-H hydroxylation reactions and, in principle, provide an attractive solution for achieving selective oxidative transformations. In practice, however, the difficulty of identifying a biocatalyst suited to a specific synthetic target often renders chemoenzymatic strategies impractical. For example, biosynthetic enzymes such as CitB and ClaD can mediate site-selective benzylic hydroxylation, enabling access to select natural products that share structural features with their native substrates. Nonetheless, such approaches typically lack generality and do not readily extend to broader collections of natural products. Here, we explore the previously unstudied natural sequence space surrounding biosynthetic hydroxylases CitB and ClaD to generate a bespoke panel of uncharacterized enzymes that collectively expand the substrate scope of biocatalytic benzylic hydroxylation reactions. Development of this curated biocatalyst panel enabled the hydroxylation of a set of orcinaldehyde-like substrates, which were further elaborated to six natural products.
Gonzalo J Villegas Rodríguez, Evan O Romero, Jonathan C Perkins et al.· Angewandte Chemie· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.