BACKGROUND
The pedagogical depiction of eukaryotic gene structure seems to assume that coding sequences (CDSs) are predefined in the genome, with transcript diversity arising mainly from exon shuffling. However, whether such "predefined CDS" model is universal remains untested.
METHODS
We systematically analyzed seven representative eukaryotic genomes to classify protein-coding genes (PCGs) into four classes based on the positional consistency of CDS start/stop sites. Both strict and loose criteria were applied, followed by cross-species comparisons of genomic feature and functional enrichment.
RESULTS
Predefined CDS genes (Class 1) were unexpectedly rare, comprising < 10% of PCGs in most species but exceeding 25% in Drosophila. Class 1 genes displayed more exons, longer CDSs, but minimal splicing isoforms, indicating purifying selection on molecular diversity. Class 1 genes are enriched in housekeeping terms like neuronal and developmental processes, whereas highly variable Class 4 genes (with distinct CDS start/end positions across different transcripts) are associated with fast-evolving processes like metabolism and reproduction.
CONCLUSIONS
In contrast to the pedagogical simplification, our results show that CDSs are rarely predefined in the genome. The differential roles of predefined versus variable CDS architectures may reflect how natural selection balances molecular stability and functional innovation.
Qian Cao, Ziyi Wang, Y. Duan· BMC Genomics· 0 citations
Abstract Both genomic mutations and RNA editing contribute to functional complexity and drive adaptive evolution. Single-cell profiling offers deep insight into the cis-regulatory mechanisms underlying these variations. Using 13 025 single-cell Smart-Seq libraries from whole-body Drosophila melanogaster, we unexpectedly found that 94.0% of adenosine-to-inosine RNA editing sites and 92.8% of heterozygous single nucleotide polymorphismss (SNPs) with sufficient “unique fragment support” exhibit binary expression (0 or 1) in a single cell. The genotypes of representative heterozygous SNPs were validated by Sanger sequencing. Meanwhile, binary RNA editing itself is logically questionable due to elusive mechanism, compromised condition specificity, and untenable heterozygote advantage. This fact that for most cases in Smart-Seq, only a single allele (out of the various haplotypes) is finally maintained per cell, raises the following concern. Regardless of the biological or technical explanations like monoallelic transcriptional burst, dropout, or amplification bias that might account for this binary expression pattern, our findings conservatively indicate that Smart-Seq may not be good at analyzing molecular diversity and that the results need to be interpreted with caution.
Y. Duan, Jiyao Liu, Shiwen Xu et al.· Nucleic Acids Research· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.