Raygun is introduced, a generative artificial intelligence framework that enables miniaturization, modification and augmentation of proteins, using a probabilistic encoding of protein sequences constructed from language model embeddings, enabling the kind of coordinated, large-scale sequence modifications that characterize natural protein evolution.
Abstract
Proteins have evolved over billions of years through coordinated substitutions, insertions and deletions, yet computational protein design cannot fully replicate nature's ability to engineer new proteins from existing templates. Protein language models1-3 generate informative per-residue representations, but harnessing them for large-scale, function-preserving sequence modifications has remained beyond reach. Here we introduce Raygun, a generative artificial intelligence framework that enables miniaturization, modification and augmentation of proteins, using a probabilistic encoding of protein sequences constructed from language model embeddings. Our key conceptual advance is to encode each protein not as a sequence of variable length in high-dimensional space, but as a probability distribution in fixed dimensions, making proteins of any length directly commensurable. Controlled by just two parameters governing substitutions and length changes, Raygun can shrink proteins by 10-25% (sometimes more than 50%), expand them beyond their natural size, and introduce extensive sequence diversity, all while preserving predicted structural integrity and functional sites. In cell-based validation, Raygun miniaturized fluorescent proteins (2 shorter than 96% of fluorescent proteins in FPbase) and TurboID, a synthetic biotin ligase that has been widely adopted for proteomics. It also expanded epidermal growth factor (EGF), generating variants with higher EGFR-binding affinity than the wild type. These results show that protein function can be faithfully captured in a length-agnostic representation, enabling the kind of coordinated, large-scale sequence modifications that characterize natural protein evolution.
An approach to create novel, functional proteins through the integration of deep mutational scanning, structural analysis, and evolutionary mining within prompts for a generative protein language model (PLM) is described and the utility of this approach is demonstrated with the generation of novel compact RNA-guided nucleases.
Nicholas W. Hughes, Sourab Kulkarni, Grant Goldman et al.· bioRxiv· 0 citations
How solid-phase peptide synthesis, genetically encoded libraries, and high-throughput selection and screening enable systematic exploration of vast, noncanonical landscapes largely inaccessible to traditional engineering is discussed.
Filip Buchel, V. G. Giacobelli, K. Hlouchová· TIBS -Trends in Biochemical...· 0 citations
MULTI-evolve is a model guided, universal, targeted installation of multimutants framework that rapidly designs hyperactive multimutant proteins and improves the identi fi cation of productive mutations compared with individual PLMs alone.
J. Koo, Young-Ho Park, Sun-Uk Kim· Signal Transduction and Targ...· 0 citations
Protein sequence space is vast due to the combinatorial diversity of 20 amino acids. However, evolution has generated a limited set of “old” canonical protein families sharing evolutionary ancestry, structures and functions. It remains unclear how canonical sequences are placed in sequence space, how recently evolved “young” proteins compare to them, and whether random, young, and canonical sequences can interconvert along evolutionarily plausible paths, and which biophysical properties distinguish or link these sequences. Here, we analyse naturally occurring de novo proteins from yeast and flies, which originate from non-coding DNA and thus have experienced limited evolutionary selection. They serve as a model for examining the relationships between young de novo and intergenic proteins, older canonical proteins, and their randomized counterparts. Because de novo and randomized sequences lack detectable homology, we use an alignment-free k-mer-based distance approach. Randomization shifts distance distributions toward expected random behaviour in all classes, but natural, non-randomized sequence classes remain distinct, indicating non-random residue organization. Each class exhibits characteristic k-mer patterns, with de novo proteins clearly separated from both canonical and all randomized sequences. Sequences bridging these classes are frequently predicted to contain transmembrane helices. De novo proteins are thus not random samples of sequence space. Instead, they occupy constrained yet evolutionarily accessible regions defined by residue order and biophysical constraints, suggesting a plausible pathway for the emergence and diversification of new proteins. Significance Statement Despite the vast combinatorial potential of amino acids, evolution has produced only a limited repertoire of canonical proteins with conserved structure and function. How evolutionarily young proteins relate to older canonical proteins, and whether the sequence space between them is traversable, remain unclear. Here, we decompose canonical proteins, intergenic sequences, and recently emerged yeast and fly de novo proteins, together with randomized controls, into short, interpretable fragments (k-mers) and compare them using alignment-free distances. De novo proteins are markedly distinct from both randomized and canonical sequences. Notwithstanding their evolutionary distance, sequences are connected by stepwise paths comprising bridge sequences, often enriched for low-complexity motifs and transmembrane helices, connecting disordered and structured regions of sequence space.
Lars A. Eicholt, Á. Tóth-Petróczy, R. Goldstein et al.· bioRxiv· 1 citation
Evolution guides biological systems to populate ecological niches, with viruses among the most successful examples of this principle. Viruses evolved over billions of years to efficiently transfer genetic information. Although viruses are highly diverse, most have converged towards remarkable similarity in the size and shape of their capsids1,2. By contrast, generative models for protein design enable the creation of protein architectures that are absent from nature3-5. Here we investigate whether protein assemblies designed by artificial intelligence can be functionalized to construct nucleic acid transport vehicles that are independent of evolutionary trajectories. By combining natural protein domains with synthetic protein assemblies, we create more than 100 bottom-up RNA transfer vehicles with unique sizes and shapes. These vehicles surpass the RNA transfer efficiency of widely used delivery vehicles by several orders of magnitude. In addition, we demonstrate that their tropism can be programmed by incorporation of computationally designed peptide binders and use them to deliver therapeutically relevant cargo RNAs into a wide range of cellular models. We show the in vivo biodistribution of one of these vehicles in a mouse at near-single-cell resolution, confirm its safety, and use it to perform a gene-editing treatment strategy for Duchenne muscular dystrophy in patient-derived cells and a pig. Our work demonstrates how proteins created by generative artificial intelligence can be harnessed for the rational engineering of RNA transport systems with the desired properties by overcoming the limitations of natural protein diversity.
Maren Kirstin Schuhmacher, Christoph Gruber, Christopher M. R. Lang et al.· Nature· 1 citation
The answer lies in geometry: proteins with denser cores, larger size, and higher-order oligomeric assembly tolerate mutations more readily, occupy larger structural families, and support more versatile biological roles, reveals that protein size, shape, and self-assembly, not just sequence, are fundamental drivers of evolvability.
S. Acharya, Sucharita Dey· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.