A language-model-guided framework that iteratively refined industry-optimized coding sequences of clinical-stage therapeutics through synonymous exploration of codon space establishes directed evolution as a practical strategy to improve biologic expression, a key manufacturing bottleneck, without altering protein sequence.
Abstract
Directed evolution is commonly used in protein engineering, where mature molecules are routinely improved through iterative local search of amino acid space. Here, we extend this principle to coding DNA. We developed a language-model-guided framework that iteratively refined industry-optimized coding sequences of clinical-stage therapeutics through synonymous exploration of codon space. Across 23 antibody-based therapeutics, SynCodonLM-guided refinement significantly increased recombinant expression in CHO cells for 17 molecules (74% responder rate), without significant compromise of product-quality or biophysical attributes. Moreover, changes in model likelihood predicted expression gains more effectively than heuristic statistical or mRNA-structure descriptors, despite no explicit expression objective. Codon-level likelihood also tracked temporal progression in influenza A H1N1 sequences, indicating the model captures evolutionary signal. These results show that even production-optimized sequences retain accessible fitness in synonymous codon space, establishing directed evolution as a practical strategy to improve biologic expression, a key manufacturing bottleneck, without altering protein sequence.
Codon optimization uses synonymous sequence changes to improve the expression and therapeutic performance of nucleic acid-based medicines. Masked language models (MLMs) have recently been proposed as alternatives to traditional, frequency-based codon optimization approaches, yet whether they offer a meaningful advantage over such simpler methods remains unclear. Here we benchmark three prominent MLMs – CaLM, EnCodon and CodonTransformer – across backtranslation fidelity, sequence generation and nine molecular phenotype prediction tasks, and experimentally evaluate model-designed sequences using a secreted embryonic alkaline phosphatase (SEAP) reporter. The models differed markedly in amino-acid fidelity and generated distinct synonymous sequence variants. However, no single model performed best across all benchmark tasks and simple sequence features remained competitive in several settings. Our interpretability analysis revealed that the models integrate a large window of codon context for making predictions, as opposed to frequency-based approaches. Our in vitro data showed that MLM-designed variants outperformed conventional and commercial-vendor-derived sequences in both transient and stably integrated expression, supporting the models’ ability to capture translational context beyond codon frequency. Together, our results establish MLMs as effective and complementary tools for codon optimization and suggest that sampling across multiple models may improve the likelihood of identifying high-performing therapeutic sequences.
Shushan Toneyan, Kerstin Scholz, Carlo De Donno et al.· bioRxiv· 0 citations
MULTI-evolve is a model guided, universal, targeted installation of multimutants framework that rapidly designs hyperactive multimutant proteins and improves the identi fi cation of productive mutations compared with individual PLMs alone.
J. Koo, Young-Ho Park, Sun-Uk Kim· Signal Transduction and Targ...· 0 citations
In this review, a review of recent in vivo hypermutation tools that enable rapid sampling of the vast evolutionary landscape, all while supporting simultaneous selection of the best proteins within living organisms are discussed.
An approach to create novel, functional proteins through the integration of deep mutational scanning, structural analysis, and evolutionary mining within prompts for a generative protein language model (PLM) is described and the utility of this approach is demonstrated with the generation of novel compact RNA-guided nucleases.
Nicholas W. Hughes, Sourab Kulkarni, Grant Goldman et al.· bioRxiv· 0 citations
Designing effective gene and mRNA sequences is a difficult optimisation problem because the number of possible nucleotide combinations grows extremely quickly with sequence length. Traditional optimisation methods such as simulated annealing are well suited to exploring these large search spaces, but their performance depends heavily on the quality of the scoring function used to evaluate candidate sequences. Hand crafted scoring rules are often slow to compute and cannot easily adapt to different biological contexts or patient specific constraints. This project presents an AI-guided simulated annealing framework for automated gene sequence design using two approaches. The first replaces fixed rule-based scoring with an adaptive model evaluating candidates using biological reference data and patient-specific information. By adjusting biological trait importance based on age, disease background, and treatment goals, the scoring model dynamically changes sequence evaluation without modifying the optimization algorithm. The second approach employs Gradient Boosting Regression on CRISPR guide RNA sequences with extracted biological features including GC content, positional nucleotides, and sequence complexity metrics. This model learns from validated literature guides, providing interpretable, deterministic scoring while maintaining adaptability. The framework is designed to support long running and repeated simulated annealing searches with minimal human intervention. Sequence evaluation is decoupled from the optimisation engine so that scoring models and reference databases can be updated as new experimental or clinical data becomes available. This allows the same optimisation pipeline to be reused across different applications such as vaccine design, cancer related gene targets or personalised therapies. By combining a fast native optimisation core with an adaptive and context aware evaluation model, this work demonstrates a flexible approach to large scale gene sequence optimisation. The proposed system highlights how AI driven scoring can improve the practicality of heuristic search methods and move sequence design closer to personalised and data driven biomedical applications.
Dai Duong Nguyen, Ryan Shaw, Kenneth Y T Lim· AHFE International· 0 citations
Abstract Standard probabilistic models of coding sequence evolution effectively identify where and when selection acts but remain agnostic to the mechanistic realization of these forces. We introduce PRIME (PRoperty Informed Models of Evolution), a framework of codon-level maximum likelihood methods—including global (G-PRIME), episodic (E-PRIME), and site-specific (S-PRIME) implementations—that explicitly model amino acid exchangeability as a function of physicochemical properties. By parameterizing attributes such as molecular volume, hydropathy, and secondary structure propensities, PRIME aims to resolve the biophysical basis of selective constraint across both the sequence and the phylogeny. At the site level, S-PRIME leverages an explicit biophysical taxonomy to categorize residues as conserved, neutral, or changing for specific properties, resolving selective signals that are missed by traditional rate-based metrics. Our analysis of a benchmark of 24 diverse datasets and a genome-wide screen of 18,944 mammalian genes demonstrates that consideration of biophysical realism can yield substantial improvements in model fit, acting synergistically with rate variation to explain complex evolutionary patterns. We find that physicochemical constraints at individual sites can be reliably detected in datasets with sufficient information redundancy (substitutions per unique amino acid; AUC=0.91), with sensitivity exceeding 90% in data-rich alignments. E-PRIME reveals a distinct hierarchy in biophysical constraints: while core packing and beta-sheet scaffolds are rigidly conserved, alpha-helix propensity and surface electrostatics serve as the primary substrates for adaptive tuning. Furthermore, PRIME importance weights align with aspects of the primary semantic axes of deep learning representations (ESM-2) and capture key features of experimental fitness landscapes. By transforming abstract evolutionary rates into interpretable biophysical rules, PRIME provides a useful framework for characterizing the mechanistic drivers of protein diversity.
Hannah Kim, Konrad Scheffler, Anton Nekrutenko et al.· Molecular biology and evolut...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.