Skip to content
Open access

Protein Structure Characters in the Light of Phylogenetic Systematics

Jul 2026 · Genome Biology and Evolution · Vol 18 · 0 citations · 25 references
Medicine

TL;DR

It is concluded that differences in 3Di characters between semaphoronts are not intrinsically a problem, but they do require that the researcher uses the same replicable method on all proteins in the phylogenetic analysis.

Abstract

Abstract Protein structure characters have great potential for improving phylogenetic inference, especially for deep nodes where amino acid sequences are highly diverged. The combination of AlphaFold structure predictions and Foldseek's “3Di” structural alphabet makes it relatively easy to conduct model-based phylogenetic inference that includes a partition of slow-evolving 3Di characters. However, we show that even identical amino acid sequences can produce substantially different 3Di characters, depending on the source of the structural model and whether inter-chain interactions are considered. We argue that such variability can be addressed with key concepts from traditional organism-based phylogenetic systematics: semaphoront, hypodigm, and character ascertainment method. To illustrate this, we develop an analogy between organismal development, taphonomy, and subsequent description and character coding by a systematist, and the process of protein synthesis, folding, and interaction and subsequent extraction, experimentation, and structural modeling by a biochemist. We conclude that differences in 3Di characters between semaphoronts are not intrinsically a problem, but they do require that the researcher uses the same replicable method on all proteins in the phylogenetic analysis. The guiding principle should be to maximize the chance that character differences in the data matrix are the results of underlying evolutionary changes, rather than artifacts due to differences in the methods used for obtaining semaphoronts and coding characters.

Read PDF

Similar papers

Open access Aug 2026

The physicochemical basis of protein evolution: property-informed evolutionary models (PRIME)

Abstract Standard probabilistic models of coding sequence evolution effectively identify where and when selection acts but remain agnostic to the mechanistic realization of these forces. We introduce PRIME (PRoperty Informed Models of Evolution), a framework of codon-level maximum likelihood methods—including global (G-PRIME), episodic (E-PRIME), and site-specific (S-PRIME) implementations—that explicitly model amino acid exchangeability as a function of physicochemical properties. By parameterizing attributes such as molecular volume, hydropathy, and secondary structure propensities, PRIME aims to resolve the biophysical basis of selective constraint across both the sequence and the phylogeny. At the site level, S-PRIME leverages an explicit biophysical taxonomy to categorize residues as conserved, neutral, or changing for specific properties, resolving selective signals that are missed by traditional rate-based metrics. Our analysis of a benchmark of 24 diverse datasets and a genome-wide screen of 18,944 mammalian genes demonstrates that consideration of biophysical realism can yield substantial improvements in model fit, acting synergistically with rate variation to explain complex evolutionary patterns. We find that physicochemical constraints at individual sites can be reliably detected in datasets with sufficient information redundancy (substitutions per unique amino acid; AUC=0.91), with sensitivity exceeding 90% in data-rich alignments. E-PRIME reveals a distinct hierarchy in biophysical constraints: while core packing and beta-sheet scaffolds are rigidly conserved, alpha-helix propensity and surface electrostatics serve as the primary substrates for adaptive tuning. Furthermore, PRIME importance weights align with aspects of the primary semantic axes of deep learning representations (ESM-2) and capture key features of experimental fitness landscapes. By transforming abstract evolutionary rates into interpretable biophysical rules, PRIME provides a useful framework for characterizing the mechanistic drivers of protein diversity.

Hannah Kim, Konrad Scheffler, Anton Nekrutenko et al. · 0 citations
Open access Jul 2026

How are evolutionarily young and old proteins distributed in sequence space?

Protein sequence space is vast due to the combinatorial diversity of 20 amino acids. However, evolution has generated a limited set of “old” canonical protein families sharing evolutionary ancestry, structures and functions. It remains unclear how canonical sequences are placed in sequence space, how recently evolved “young” proteins compare to them, and whether random, young, and canonical sequences can interconvert along evolutionarily plausible paths, and which biophysical properties distinguish or link these sequences. Here, we analyse naturally occurring de novo proteins from yeast and flies, which originate from non-coding DNA and thus have experienced limited evolutionary selection. They serve as a model for examining the relationships between young de novo and intergenic proteins, older canonical proteins, and their randomized counterparts. Because de novo and randomized sequences lack detectable homology, we use an alignment-free k-mer-based distance approach. Randomization shifts distance distributions toward expected random behaviour in all classes, but natural, non-randomized sequence classes remain distinct, indicating non-random residue organization. Each class exhibits characteristic k-mer patterns, with de novo proteins clearly separated from both canonical and all randomized sequences. Sequences bridging these classes are frequently predicted to contain transmembrane helices. De novo proteins are thus not random samples of sequence space. Instead, they occupy constrained yet evolutionarily accessible regions defined by residue order and biophysical constraints, suggesting a plausible pathway for the emergence and diversification of new proteins. Significance Statement Despite the vast combinatorial potential of amino acids, evolution has produced only a limited repertoire of canonical proteins with conserved structure and function. How evolutionarily young proteins relate to older canonical proteins, and whether the sequence space between them is traversable, remain unclear. Here, we decompose canonical proteins, intergenic sequences, and recently emerged yeast and fly de novo proteins, together with randomized controls, into short, interpretable fragments (k-mers) and compare them using alignment-free distances. De novo proteins are markedly distinct from both randomized and canonical sequences. Notwithstanding their evolutionary distance, sequences are connected by stepwise paths comprising bridge sequences, often enriched for low-complexity motifs and transmembrane helices, connecting disordered and structured regions of sequence space.

Lars A. Eicholt, Á. Tóth-Petróczy, R. Goldstein et al. · 1 citation
Preprint Jul 2026

Structural Compression for Phylogenetic Inference under Alignment Instability and Indel-Rich Evolution

Phylogenetic inference traditionally relies on aligned characters under substitution models, but this framework becomes less reliable when alignments are unstable or when evolution is dominated by insertions, deletions, repeats, and other structural changes. We adapt Ladderpath as an alignment-free distance approach for phylogenetic inference. Motivated by algorithmic information theory, Ladderpath decomposes sequences into derived, reusable units (``ladderons'', rather than fixed-length $k$-mers) organized hierarchically, from which pairwise distances are computed. The premise is that shared derived sequence structure, including repeated or reused segments that are poorly represented by column-wise substitutions, can retain phylogenetic information. The bacteriophage T7 known lineage, the cpSSR repeat-rich marker, and a cytochrome~$c$ protein dataset confirm that Ladderpath recovers topologies consistent with the known experimental history or with established alignment-based methods. Its advantage emerges under stress: in block-translocation and indel-dominated simulations Ladderpath remains stable while alignment-dependent pipelines deteriorate; on banana mitochondrial and plastome genomes it scales to genome length and captures the expected contrast between organellar histories, all from unaligned input. These results support Ladderpath as an alignment-free, structurally informed method that could complement standard pipelines in cases where higher-order sequence structure carries phylogenetic signal.

Zhu Zhang, Jieyu Wang, Fengyao Zhai et al. · 0 citations
Open access Jul 2026

4MRNA: a new approach for nucleic acid molecular replacement using models with diverse parameter patterns.

Structural analysis of nucleic acids lags behind that of proteins, partly because most fundamental structural analysis techniques have been primarily developed for proteins. The molecular replacement (MR) method, commonly used for phase determination in protein crystallography, encounters unique challenges when applied to nucleic acids. Nucleic acids can have different three-dimensional structures even with the same sequence, which often renders database entries or predicted models unsuitable as search models for MR. To overcome the limitation, we developed a novel strategy termed 4MRNA, which stands for Massive Multi-type Model Molecular Replacement for Nucleic Acids. This method introduces a new principle for MR, which is the systematic creation of diverse search models through parameter adjustment. By identifying the parameters that critically influence MR and generating models based on their statistical analysis, 4MRNA can provide search models that closely approximate target structures and thereby improve the success rate of MR. Its effectiveness was validated across comprehensive test cases including canonical duplexes, duplexes with bulges and internal loops, the more complex structure of transfer RNA, and a previously unreported DNA structure. 4MRNA is anticipated to become an indispensable tool for nucleic acid structure determination, profoundly advancing fundamental research and extending its impact to wide-ranging applications including structure-based drug design and nucleic acid nanotechnology.

Shin Ando, Jiro Kondo · 0 citations
#protein folding Open access Aug 2026

Evolutionary origins of protein novelty across an entire yeast subphylum

Genes encoding novel protein sequences are a ubiquitous feature of genomes. They fuel molecular and cellular evolutionary innovations and frequently contribute to species-specific characteristics. We are now unravelling the processes by which they originate, including de novo from noncoding sequences and through extreme divergence, yet how much and what types of novel proteins evolve through each process is still unclear Does the mechanism of origination shape the structural and functional potential of the resulting proteins? Here, we conducted a broad computational investigation of genetic and protein novelty at the scale of the entire subphylum of Saccharomycotina yeasts. We detected more than 5,000 robust de novo genes across 332 species and compared them to more than 10,000 novel genes resulting from extreme sequence divergence, revealing two distinct modes of evolution of novelty. A remarkable 40% of de novo proteins are predicted to localize to mitochondria compared to only 15% of divergent, with the latter also being substantially longer and more disordered. A detailed analysis of conservatively predicted tertiary structures of novel proteins shows that “invention” of novel folds can happen through both processes but is more likely to occur de novo. We also illustrate cases of evolutionary “re-invention” of existing protein folds from non-coding sequences. Our work deepens our understanding of the origins and importance of novel proteins opening new directions for further structural and functional characterization.

Emilios Tassios, Nikolaos Pirgelis, David Rinker et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.