Abstract Background Biological discovery and design are increasingly guided by predictive models trained on data from high-throughput technologies rather than costly experiments. However, existing datasets are often biased by overrepresentation of model organisms, causing models to fail in evolutionary studies of non-model species. We focus on transcriptional activators, which contain activation domains (ADs) that promote gene expression. ADs are intrinsically disordered and poorly conserved, limiting their study using comparative genomics. Results We present a hybrid framework that leverages high-throughput molecular assays and active learning to quantify biological properties across evolutionary space. We develop ADhunter, a high-capacity regression model that outperforms state-of-the-art algorithms in identifying transcriptional activators and quantifying their strength. We use model-based uncertainty to guide evolutionary sampling across 7,842,516 proteins from 2,400 fungal genomes. We functionally characterize 9,836 ADs from 1,071 fungal genomes, providing a 15.5-fold expansion in genome representation compared with existing datasets. Comprehensive sampling improves model generalizability and provides the first functional annotation for 3,416 proteins in non-model fungi. Interpretability analysis of ADhunter aligns with biophysical models and reveals novel, underrepresented protein codes. Conclusions These results highlight the importance of sampling from non-model organisms to build evolutionarily robust functional genomics models. Our framework provides a general strategy for building predictive models that better capture the diversity of natural sequence-to-function relationships.
Lucas Waldburger, Hunter Nisonoff, Marissa Zintel et al.· Genome biology· 0 citations
Ribulose-1,5-bisphosphate carboxylase/oxygenase (Rubisco) fixes the majority of carbon dioxide globally but is challenged with low specificity for CO2 versus O2 and low catalytic efficiencies. Traditional engineering efforts have remained difficult because folding, assembly, specificity, and catalysis are tightly coupled, hampering efforts to explore sequence space. Therefore, we leveraged recent advances in protein large language models (PLMs) to generate sequences beyond those observed in nature, using both ProGen-2 that was fine-tuned on a limited dataset of non-Form I Rubiscos and an ESM-2 discriminator. With this approach, we generated 5.6 million novel Rubisco-like sequences and identified 21 highly diverse candidates predicted to be active that occupy regions of Rubisco phylogenetic space not previously observed in nature. Six designs were soluble in Escherichia coli, and five were shown to produce quantifiable 3PGA. One design produced an apparent CO₂/O₂ specificity estimate beyond the range of the natural representative Rubiscos assayed. We also solved the crystal structure of one de novo design that reproduced the predicted dimer and active-site geometry with sub-angstrom Cα agreement. Sequence-only generation followed by independent structural filtering therefore recovered soluble, active Rubiscos from regions of sequence space that are not represented in genomic databases. Together, these results establish a scalable strategy for accessing previously unexplored Rubisco sequence space, providing a broadly accessible path toward generating de novo Rubiscos that may have activity and specificity parameters needed to address longstanding limitations in biological carbon fixation.
Alexander J. Kehl, Simon K. S. Chu, J.H. Pereira et al.· bioRxiv (Cold Spring Harbor...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.