Skip to content
Open access

Coding regions are rarely predefined in the eukaryotic genomes: a note on simplified models of gene architecture.

Jul 2026 · BMC Genomics · 0 citations
Medicine

Abstract

Background

The pedagogical depiction of eukaryotic gene structure seems to assume that coding sequences (CDSs) are predefined in the genome, with transcript diversity arising mainly from exon shuffling. However, whether such "predefined CDS" model is universal remains untested.

Methods

We systematically analyzed seven representative eukaryotic genomes to classify protein-coding genes (PCGs) into four classes based on the positional consistency of CDS start/stop sites. Both strict and loose criteria were applied, followed by cross-species comparisons of genomic feature and functional enrichment.

Results

Predefined CDS genes (Class 1) were unexpectedly rare, comprising < 10% of PCGs in most species but exceeding 25% in Drosophila. Class 1 genes displayed more exons, longer CDSs, but minimal splicing isoforms, indicating purifying selection on molecular diversity. Class 1 genes are enriched in housekeeping terms like neuronal and developmental processes, whereas highly variable Class 4 genes (with distinct CDS start/end positions across different transcripts) are associated with fast-evolving processes like metabolism and reproduction.

Conclusions

In contrast to the pedagogical simplification, our results show that CDSs are rarely predefined in the genome. The differential roles of predefined versus variable CDS architectures may reflect how natural selection balances molecular stability and functional innovation.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.