Skip to content
Open access

Unobserved Sequence Space Has Many Functional Proteins

Aug 2026 · bioRxiv · 0 citations · 51 references
Biology

TL;DR

It is determined that portions of protein sequence space, despite being unobserved in nature, contains many functional proteins that cannot be predicted accurately in silico, providing experimental support for the hypothesis that natural protein sequences explored by evolution represent a miniscule fraction of all possible functional sequences.

Read PDF

Similar papers

Open access Jul 2026

How are evolutionarily young and old proteins distributed in sequence space?

Protein sequence space is vast due to the combinatorial diversity of 20 amino acids. However, evolution has generated a limited set of “old” canonical protein families sharing evolutionary ancestry, structures and functions. It remains unclear how canonical sequences are placed in sequence space, how recently evolved “young” proteins compare to them, and whether random, young, and canonical sequences can interconvert along evolutionarily plausible paths, and which biophysical properties distinguish or link these sequences. Here, we analyse naturally occurring de novo proteins from yeast and flies, which originate from non-coding DNA and thus have experienced limited evolutionary selection. They serve as a model for examining the relationships between young de novo and intergenic proteins, older canonical proteins, and their randomized counterparts. Because de novo and randomized sequences lack detectable homology, we use an alignment-free k-mer-based distance approach. Randomization shifts distance distributions toward expected random behaviour in all classes, but natural, non-randomized sequence classes remain distinct, indicating non-random residue organization. Each class exhibits characteristic k-mer patterns, with de novo proteins clearly separated from both canonical and all randomized sequences. Sequences bridging these classes are frequently predicted to contain transmembrane helices. De novo proteins are thus not random samples of sequence space. Instead, they occupy constrained yet evolutionarily accessible regions defined by residue order and biophysical constraints, suggesting a plausible pathway for the emergence and diversification of new proteins. Significance Statement Despite the vast combinatorial potential of amino acids, evolution has produced only a limited repertoire of canonical proteins with conserved structure and function. How evolutionarily young proteins relate to older canonical proteins, and whether the sequence space between them is traversable, remain unclear. Here, we decompose canonical proteins, intergenic sequences, and recently emerged yeast and fly de novo proteins, together with randomized controls, into short, interpretable fragments (k-mers) and compare them using alignment-free distances. De novo proteins are markedly distinct from both randomized and canonical sequences. Notwithstanding their evolutionary distance, sequences are connected by stepwise paths comprising bridge sequences, often enriched for low-complexity motifs and transmembrane helices, connecting disordered and structured regions of sequence space.

Lars A. Eicholt, Á. Tóth-Petróczy, R. Goldstein et al. · 1 citation
Open access Jul 2026

Expanding the scope of protein language modeling to protein-protein interactions with MSA Pairformer.

Multiple sequence alignment (MSA) Pairformer is presented, a protein language model that builds on AlphaFold2/3's bidirectional refinement between sequence and pairwise residue representations to accurately model the evolution of protein-protein interactions, despite training exclusively on individual chains.

Yo Akiyama, Zhidian Zhang, Olivia Tang et al. · 3 citations
Open access Aug 2026

PLMView: collaborative protein language model representations for fast and scalable specialized protein function inference

Applications to thioredoxins, visual opsins, and Tara Oceans environmental diatom cold-shock proteins show that PLMView can move from interpretable residue-level determinants in well-studied protein families to large-scale environmental functional discovery, linking molecular specialization to ecological distribution and transcriptional deployment across the global ocean.

Vinh-Son Pho, Alessandro Natale Bianchi, Mattéo Scarsini et al. · 0 citations
Open access Aug 2026

A discrete protein subset drives structure prediction discordance in orphan proteins

The discordance persists: pLDDT correlates positively with PUNCH2 disorder in random and de novo proteins and negatively with β-strand fraction, opposite to the conserved and disordered baselines, a concrete failure mode that protein designers and other working on sequences remote in sequence space should be aware of when relying on predictor outputs.

Lars A. Eicholt, Lasse Middendorf · 0 citations
Open access Aug 2026

Large-scale structure prediction of DUF-containing protein-protein interactions

Whether AlphaFold 3 complex prediction, combined with STRING evidence and domain-level analysis of interfaces and interaction partners, can help identify and characterize DUF-containing proteins and suggest roles for DUF4130 in nucleic-acid-associated radical-SAM biology and DUF5819 in a bacterial system related to vitamin-K-dependent carboxylation are suggested.

Lino Riepenhausen, Francesco Costa, Antonina Andreeva et al. · 0 citations
Open access Aug 2026

Protein language models and the long tail of functional diversity

It is found that many GigaRef singletons belong to a cluster under alternative parameter settings, suggesting that genomic and metagenomic datasets may require dataset-specific clustering configurations, and it is shown that singletons share mutual information with clustered sequences, making them learnable by PLMs and useful for training.

R. Vinod, Samir Char, Ava A. Amini et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.