Skip to content
Open access

Pangenome discovery and characterization of human protein-coding duplicated genes

Aug 2026 · bioRxiv · 0 citations · 7 references
Biology

TL;DR

The pangenome provides unparalleled specificity to understand genetic variation in SD genes allowing us to distinguish functional genes from pseudogenes and highlighting potential gene innovations that arose most recently in human evolution.

Abstract

Protein-coding genes mapping to high-identity segmental duplications (SDs) have been difficult to annotate and characterize and are the source of most previously unknown protein-coding genes being discovered as part of the human pangenome. Here, we combine long-read assembled human genomes (298) and long-read transcriptome data (5.6 billion full-length cDNA from 83 tissues) to phylogenetically interrogate 493 gene families discovering 2713 potentially copy number polymorphic genes not present in the human reference genome. For reference SD gene families where paralog specificity can be assigned, we find that 60.0% are expressed and maintain open reading frames, with 45.7% showing high expression in brain, embryo, or testis. We revise 386 gene models, including 150 that absent or different from current T2T-CHM13 gene annotation and 236 (35.1%) pseudogenes as protein-coding where we find evidence of transcription, an open reading frame, and chromatin-accessible promoters. We find that 24.2% of SD genes show evidence of constraint for both copy number and amino acid mutation. The majority of these constraint genes are ancestral, whereas only 16.2% of derived duplicated genes that emerged recently in the human lineage show evidence of constraint. The pangenome provides unparalleled specificity to understand genetic variation in SD genes allowing us to distinguish functional genes from pseudogenes and highlighting potential gene innovations that arose most recently in human evolution.

Read PDF

Similar papers

Open access Aug 2026

Analysis of spliceosome-related coding and noncoding genes and pseudogenes reveals novel candidates

Splicing is a complex molecular mechanism in eukaryotic cells essential to gene expression and regulation, involving more than 300 protein-coding genes (PCGs) and 43 small nuclear RNA (snRNA) genes. However, fewer than 30 gene-disease relationships have been described as spliceosomopathies to date. This discrepancy suggests the splicing machinery as an underexplored area for human disease gene discovery. For snRNA currently classified as pseudogenes, we prioritized candidates with similar epigenomic, genomic, and hypermutability features as functional snRNA genes. Population-variant-depletion analysis was performed to identify regions under negative selection. We analyzed rare variants in PCGs and snRNA genes and prioritized snRNA pseudogenes across a large heterogeneous rare disease cohort. There was high concordance for prioritizing genes annotated as pseudogenes by the variant-depleted region analysis (9) and by random forest models of hypermutation, genomic and epigenomic features (6). We identified 26 variants of interest across six PCGs with established gene-disease relationships (GDRs) and 14 genes not yet disease-associated, including one pseudogene across 30 individuals. For snRNAs genes, we identified 49 variants of interest located in seven genes with established GDR and 11 genes not yet disease-associated, including two pseudogenes across 80 individuals. This study highlights the importance of splicing-related PCG and snRNA in the genetic etiology of rare diseases. By leveraging specialized approaches for prioritizing pseudogenes, combined with the PCG and snRNA analysis, the genes and variants expand the variant pathogenicity spectrum of spliceosomopathies and suggest variants for follow-up case series and future functional validation.

O. Messaoud, S. DiTroia, R. Tarawneh et al. · 0 citations
Open access Jul 2026

Coding regions are rarely predefined in the eukaryotic genomes: a note on simplified models of gene architecture.

BACKGROUND The pedagogical depiction of eukaryotic gene structure seems to assume that coding sequences (CDSs) are predefined in the genome, with transcript diversity arising mainly from exon shuffling. However, whether such "predefined CDS" model is universal remains untested. METHODS We systematically analyzed seven representative eukaryotic genomes to classify protein-coding genes (PCGs) into four classes based on the positional consistency of CDS start/stop sites. Both strict and loose criteria were applied, followed by cross-species comparisons of genomic feature and functional enrichment. RESULTS Predefined CDS genes (Class 1) were unexpectedly rare, comprising < 10% of PCGs in most species but exceeding 25% in Drosophila. Class 1 genes displayed more exons, longer CDSs, but minimal splicing isoforms, indicating purifying selection on molecular diversity. Class 1 genes are enriched in housekeeping terms like neuronal and developmental processes, whereas highly variable Class 4 genes (with distinct CDS start/end positions across different transcripts) are associated with fast-evolving processes like metabolism and reproduction. CONCLUSIONS In contrast to the pedagogical simplification, our results show that CDSs are rarely predefined in the genome. The differential roles of predefined versus variable CDS architectures may reflect how natural selection balances molecular stability and functional innovation.

Qian Cao, Ziyi Wang, Y. Duan · 0 citations
Open access Sep 2026

Exploring the Effect of Whole-Genome Duplication on Salmonid LincRNA Repertoire

Long intergenic non-coding RNAs (lincRNAs) are key epigenetic regulators of genome function, yet their evolutionary dynamics following whole-genome duplication (WGD) events remain poorly understood. Salmonids, which underwent a lineage-specific autotetraploidization (salmonid-specific WGD, ~88–100 million years ago), provide an excellent model to investigate the retention, divergence, and functional potential of recently duplicated non-coding elements. LincRNA repertoires were compared across five genome-annotated salmonids (Oncorhynchus tshawytscha, O. kisutch, O. mykiss, Salmo salar, and S. trutta) and their closest non-duplicated relative, northern pike (Esox lucius). LincRNAs represented ~5–7% of annotated genes in all salmonids except S. salar (18%). Sequence conservation was low relative to coding genes, with only 11–68 highly similar (e-value < 1 × 10−30; similarity > 70% and alignments > 100 nucleotides) putative orthologues shared between salmonids and northern pike, and 161–338 among salmonids alone. Synteny conservation was modest in lincRNAs, with lower conservation in putative orthologues (8–16%) compared to putative ohnologues (8–33%). Secondary structure conservation was associated with sequence similarity (ρ = −0.45; p = 2.2 × 10−16), and the association was stronger among WGD ohnologues than orthologues. In S. salar and O. mykiss, lincRNA putative ohnologues showed weaker expression correlations than coding genes, suggesting widespread regulatory divergence, possibly through neo- and subfunctionalisation. Conserved salmonid lincRNAs showed enriched predicted interactions with miRNAs involved in tumour suppression, brain, bone, and muscle development (e.g., miR-455, miR-365, miR124, miR-133a, miR-140, and miR-9), a finding supported by limited transcriptomic data. Although salmonid WGD expanded lincRNA repertoires, lincRNAs have undergone rapid sequence and transcriptional divergence, with limited conservation across species based on sequence similarity, chromosomal position, synteny, and secondary structure. A subset of conserved lincRNAs retains structural features and regulatory signatures consistent with roles as miRNA sponges in brain, skeletal, and muscle development and tumour suppression, potentially acting within conserved regulatory networks. These findings provide new insights into lincRNA evolution following genome duplication and highlight the need for experimental validation of their regulatory functions.

Isabel García-Pérez, D. García de la Serrana · 0 citations
Review Open access Jul 2026

Hidden proteins encoded by non-canonical open reading frames: A review.

There is increasing evidence that translation is not limited to annotated protein-coding genes. Ribosome profiling sequencing, mass spectrometry-based proteomics, and immunopeptidomics have identified the productive translation of non-canonical open reading frames (ORFs). This suggests that the functional proteome includes not only conserved proteins but also proteins hidden in non-coding RNAs and de novo proteins. Some of these translated products are functional peptides, while others may be non-functional, potentially arising from evolutionary events. Several non-canonical ORF-encoded peptides have been found to regulate multiple physiological and pathological functions, particularly in cancer, immunity, and inflammation, indicating that they have potential as biomarkers and novel therapeutic targets. To better understand the diversity of functional peptides and translated non-canonical ORFs based on existing data, we summarize their classification according to transcriptional features and supporting evidence, including non-canonical ORFs located in ncRNAs and canonical mRNAs. This review provides a concise summary of the origin, discovery methods, and classification of non-canonical ORFs. It offers insights into the origins and functions of non-canonical ORF-encoded peptides from an evolutionary perspective, while also exploring the biological functions and regulatory mechanisms of these non-canonical ORF-encoded hidden proteins in tumorigenesis and progression.

Jiang-Xue Li, Yanxia Qin, Zhao Li et al. · 0 citations
Open access Aug 2026

A high-resolution human pangenome structural variant resource for improved disease association

This work describes the full spectrum of genetic variation and shows that while 99% of the variants between any two genomes are single base-pair substitutions, 88% of the euchromatic variant base pairs are SVs, including insertions, deletions, duplications, and inversions.

J. Lin, J. Gustafson, J. Wertz et al. · 0 citations
Open access Aug 2026

Neuronal Gene Architecture in Cancer borealis Revealed by Long-Read Genome Assembly and Deep Transcriptomic Analysis

Understanding the underlying neuronal function in non-model organisms requires accurate resolution of gene structure and transcript diversity. Here, we present a comprehensive genome annotation tor the Jonah crab (Cancer borealis), a key experimental system in crustacean neurobiology, with a particular focus on transcriptome-supported neuronal gene architecture. By integrating long-read genome assembly with extensive transcriptomic evidence, we reconstructed gene models with high confidence, enabling detailed characterization of exon-intron organization, alternative splicing, and isotorm diversity across gene families. Functional classification revealed extensive representation of neural-associated gene classes, including ion channels and receptors, transporters, enzymes, zinc finger proteins, histones, structural proteins, and cell adhesion molecules, alongside a large set of previously uncharacterized genes. In this study we particularly focused on the neuronal and ion channel gene families known to underlie circuit-level neuronal function in C. borealis. We provide an in-depth analysis of 87 genes spanning 17 neural-related gene families and 41 neuropeptides, detailing chromosomal localization, gene length, exon-intron configuration, and transcript-supported isotorm structure. For many of these genes, transcriptomic data confirmed expression and refined coding boundaries. Comparisons with existing transcriptomic datasets demonstrate strong concordance in gene expression patterns while also revealing novel transcripts and expanded gene family members not previously annotated. Together, this genome and transcriptome-integrated annotation establishes a high-resolution framework tor studying neuronal gene organization in C. borealis. T his resource enables direct connections between gene architecture, transcript diversity, and neural function, supporting future investigations in crustacean neurogenomics, comparative genomics, and the evolution of nervous system complexity.

Murugesan Raju, Adam J. Northcutt, David J. Schulz · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.