An approach to create novel, functional proteins through the integration of deep mutational scanning, structural analysis, and evolutionary mining within prompts for a generative protein language model (PLM) is described and the utility of this approach is demonstrated with the generation of novel compact RNA-guided nucleases.
EvoMax is established as an integrated strategy for engineering compact eukaryotic Fz2 genome editors and FanzMAX v3-hLa is identified as a high-efficiency programmable nuclease for mammalian genome editing.
Shijie Wan, Jackson Gold, Pranay Vure et al.· Nature Biotechnology· 0 citations
Raygun is introduced, a generative artificial intelligence framework that enables miniaturization, modification and augmentation of proteins, using a probabilistic encoding of protein sequences constructed from language model embeddings, enabling the kind of coordinated, large-scale sequence modifications that characterize natural protein evolution.
Kapil Devkota, Daichi Shonai, Joey Mao et al.· Nature· 1 citation
A protein design strategy is used that couples a structure-guided inverse-folding model with evolution-informed residue constraints to generate active, divergent variants of TnpB, a minimal CRISPR-Cas12-like nuclease, termed SynTnpBs, establishing a strategy for creating non-natural RNA-guided nucleases and conformationally active nucleic acid binders, enlarging the designable protein space.
Petr Skopintsev, Isabel Esain-Garcia, Evan C. DeTurk et al.· Science· 2 citations
In this review, a review of recent in vivo hypermutation tools that enable rapid sampling of the vast evolutionary landscape, all while supporting simultaneous selection of the best proteins within living organisms are discussed.
Protein sequence space is vast due to the combinatorial diversity of 20 amino acids. However, evolution has generated a limited set of “old” canonical protein families sharing evolutionary ancestry, structures and functions. It remains unclear how canonical sequences are placed in sequence space, how recently evolved “young” proteins compare to them, and whether random, young, and canonical sequences can interconvert along evolutionarily plausible paths, and which biophysical properties distinguish or link these sequences. Here, we analyse naturally occurring de novo proteins from yeast and flies, which originate from non-coding DNA and thus have experienced limited evolutionary selection. They serve as a model for examining the relationships between young de novo and intergenic proteins, older canonical proteins, and their randomized counterparts. Because de novo and randomized sequences lack detectable homology, we use an alignment-free k-mer-based distance approach. Randomization shifts distance distributions toward expected random behaviour in all classes, but natural, non-randomized sequence classes remain distinct, indicating non-random residue organization. Each class exhibits characteristic k-mer patterns, with de novo proteins clearly separated from both canonical and all randomized sequences. Sequences bridging these classes are frequently predicted to contain transmembrane helices. De novo proteins are thus not random samples of sequence space. Instead, they occupy constrained yet evolutionarily accessible regions defined by residue order and biophysical constraints, suggesting a plausible pathway for the emergence and diversification of new proteins. Significance Statement Despite the vast combinatorial potential of amino acids, evolution has produced only a limited repertoire of canonical proteins with conserved structure and function. How evolutionarily young proteins relate to older canonical proteins, and whether the sequence space between them is traversable, remain unclear. Here, we decompose canonical proteins, intergenic sequences, and recently emerged yeast and fly de novo proteins, together with randomized controls, into short, interpretable fragments (k-mers) and compare them using alignment-free distances. De novo proteins are markedly distinct from both randomized and canonical sequences. Notwithstanding their evolutionary distance, sequences are connected by stepwise paths comprising bridge sequences, often enriched for low-complexity motifs and transmembrane helices, connecting disordered and structured regions of sequence space.
Lars A. Eicholt, Á. Tóth-Petróczy, R. Goldstein et al.· bioRxiv· 1 citation
Protein assemblies, such as fibers, cages, and sheets, are essential components of biological systems, with versatile functions that make them attractive engineering targets for biotechnological applications. Understanding the complex sequence–structure–function relationships that govern these assemblies is critical for both basic science and the engineering of novel nanomaterials. Deep mutational scanning (DMS) has emerged as a powerful technique for mapping these relationships across large sections of protein sequence space. Specifically, DMS couples high‐throughput assays with next‐generation sequencing technologies to create datasets that report on how changes to protein sequence alter protein function. This review provides an overview of protein assemblies and the basic principles of DMS, followed by a discussion of how DMS has been applied to protein assemblies, and what unique considerations arise when performing such studies. We aim to provide a comprehensive foundation for researchers across biochemistry and chemical biology looking to leverage such high‐throughput approaches to understand and engineer the next generation of protein‐based assemblies.
Jenna B. Wolfanger, Shoili Banerjee, Carolyn E. Mills· Chemistry–Methods· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.