Skip to content
Open access

First breed-pool whole genome sequencing of egyptian sheep: a comprehensive genomic atlas revealing diversity and candidate genes for production and adaptation

Jul 2026 · BMC Biotechnology · Vol 26 · 0 citations · 82 references
Medicine

TL;DR

This research provides the first extensive breed‑pool whole‑genome sequencing (WGS) analysis across five Egyptian sheep populations, establishing a genomic atlas for the genetic architecture of production and adaptation in Egyptian sheep, providing a baseline for future genetic and conservation strategies.

Abstract

This research provides the first extensive breed‑pool whole‑genome sequencing (WGS) analysis across five Egyptian sheep populations: Barki (BAR), Rahmani (RAH), their crossbred offspring (CRS), Ossimi (OSI) and Awassi (AWI). To establish a genomic atlas of the genetic architecture of production and adaptation in Egyptian sheep, providing a baseline for future candidate gene discovery and conservation strategies. Through Illumina sequencing of 120 samples, we compiled a dataset exceeding 470 Gb, with mean coverage depths spanning 24.2x to 41.3x. Variant profiling, functional annotation, KEGG pathway analysis, and independent structural variant analysis were conducted. Phenotypic data were collected and validated through qRT-PCR gene expression analysis. Variant profiling revealed between 11.9 and 17.4 million SNPs per breed after stringent filtering. Heterozygosity patterns (population‑level estimates) differed substantially between groups, recorded at 60.41% in the CRS crossbred versus 74.92–85.01% in the purebred lines. Functional annotation identified conserved enrichment related to xenobiotic detoxification and lipid metabolism. KEGG pathway analysis prioritized the PPAR signalling pathway (map03320) and fatty acid metabolism (map01212) as highly significant (p < 0.0001). Independent structural variant analysis identified distinct genomic hotspots on chromosomes 2, 6 and 18, overlapping candidate genes; FABP4, KAP cluster and MSTN implicated in the regulation of fat deposition and muscle development. Phenotypic data confirmed a high degree of breed divergence (p < 0.001). RAH and CRS individuals reached higher body condition scores (BCS 4.31‑ 4.53) and increased fat deposition, whereas BAR was significantly leaner (BCS 2.53). The highest trimmed meat yields were observed in CRS (23.95 kg) and RAH (18.40 kg) (p < 0.001), with RAH also displaying the highest intramuscular fat content at 4.20%. qRT‑PCR validation showed elevated expression of lipogenic genes (ACACA, FASN and FABP4) in fat‑tailed breeds and differential expression of myogenic regulators (MSTN and IGF‑1) correlating with muscularity variations. The current findings establish a genomic atlas for the genetic architecture of production and adaptation in Egyptian sheep, providing a baseline for future genetic and conservation strategies. Formal selection signature analyses, such as XP‑EHH and iHS are recommended for subsequent studies.

Read PDF

Similar papers

Aug 2026

Comparative genomics of Standardbred horses reveals candidate regions under selection for harness racing performance.

BACKGROUND Harness racing performance in horses is a complex polygenic trait influenced by genetic background, breeding history, training, and environment. Understanding genomic patterns shaped by long-term artificial selection provides insights into performance-related traits and breed-specific genomic variation. AIMS/OBJECTIVES This study aimed to characterize genomic differentiation, population structure, and candidate genomic regions potentially influenced by historical selection in Standardbred horses using comparative whole-genome sequencing involving diverse horse breeds. METHODS Whole-genome sequencing data from 86 horses representing 16 breeds were retrieved from public repositories and processed using a standardized bioinformatics pipeline. Population structure was investigated using principal component analysis (PCA), ADMIXTURE, Neighbor-Joining phylogeny, and ChromoPainter-based haplotype sharing. Genomic diversity was evaluated using runs of homozygosity (ROH) and nucleotide diversity (π), whereas genetic differentiation was assessed using fixation index (FST) analyses with functional annotation. RESULTS After quality control, 18,384,176 autosomal SNPs were retained. Standardbred horses consistently formed a distinct genomic group, with principal components 1 and 2 explaining 5.00% and 3.66% of genomic variation, respectively. ADMIXTURE supported K = 2 as the best-supported clustering solution (cross-validation error = 0.523). ROH analyses revealed variation in genome-wide homozygosity among breeds, while FST and nucleotide diversity analyses identified differentiated genomic regions containing candidate genes associated with growth, skeletal development, muscle function, signaling, metabolism, and nervous system function, including LCORL, NCAPG, PDE1A, CDH13, and HTR1A. CONCLUSION This comparative whole-genome analysis identified candidate genomic regions potentially shaped by historical selection in Standardbred horses and contributes to understanding genomic differentiation associated with breed history and functional specialization.

I. Moazami, M. Mohammadabadi, H. A. Nanaei et al. · 0 citations
Open access Jul 2026

Uncovering Functional Genetic Variation in the Indigenous Greek Eghoria Goat: An Integrative Genome-Wide Discovery and Validation Approach

The Eghoria goat constitutes the major indigenous goat population in Greece and represents an important genetic resource. Despite its significance, genomic information remains scarce. This study aimed to investigate coding variation in the Eghoria goat using whole-genome sequencing (WGS), with particular focus on missense single nucleotide polymorphisms (SNPs) and coding insertions/deletions (indels). Whole-genome sequencing was performed on Eghoria goats (n = 6), followed by bioinformatic processing and variant filtering. Functional characterization was conducted through Gene Ontology (GO) and KEGG pathway analyses (FDR < 0.05). Normalized variant density metrics identified highly polymorphic genes. Selected missense SNPs were validated in an independent population (n = 54) using MassARRAY genotyping. Whole-genome analysis identified 10,796,211 SNPs and 1,022,779 indels. After prioritization, 15,949 missense SNPs and 861 indels were retained. Functional enrichment analyses highlighted pathways related to metabolism, ion transport, calcium signaling, immune function, transcriptional regulation, and environmental adaptation. Genes exhibiting elevated polymorphism density included olfactory receptor family members, and loci associated with metabolic regulation. Validation analyses confirmed the presence of selected variants in the Eghoria goat population. The study expands current knowledge of genomic diversity in the indigenous Eghoria goat and provides valuable resources for future genetic improvement, conservation, and breeding programs.

Maria-Anna Kyrgiafini, G. Stamatellos, C. Stamatis et al. · 0 citations
Open access Sep 2026

Genome-Wide Differentiation, Inbreeding, and Candidate Selection Loci in Local Vietnamese Pig Breeds

Vietnam harbors exceptional genetic diversity among at least 26 indigenous pig breeds. We analyzed genome-wide single-nucleotide polymorphism (SNP) data from 90 animals representing 15 local Vietnamese breeds and six Landrace pigs using principal component analysis, the windowed fixation index (FST), cross-population extended haplotype homozygosity (XP-EHH), within-population integrated haplotype score (iHS), and runs of homozygosity (ROHs). The population structure was consistent with a north–south differentiation axis, and Ba Xuyen showed elevated heterozygosity, providing suggestive evidence of a European genetic contribution; the f3 statistic was positive (f3 = +0.015), and formal evidence of admixture requires a significantly negative f3, so this criterion was not met. Integration of FST and XP-EHH identified GPC5, E2F6, NOS1, and TLR4 as top Northern candidate loci and CRYM/ZP2 as the leading Central candidate locus, and these windows were recovered at both the 90th and 95th percentile thresholds, indicating analytical robustness rather than independent biological validation. iHS was elevated at E2F6 in Northern breeds (|iHS| = 3.04) and at NOS1 across all regional groups (|iHS| = 2.66–3.36). Breed-level phenotypic XP-EHH, based on published breed descriptions and coat color rather than individual body-composition measurements, identified GALNT2 as a candidate shared across breed groups; HCAR1 and ATG10 as candidates specific to the extreme-fat/prolific breed group; and EFNA5 and HIPK2 as candidates specific to the medium-bodied breed group. ROHs identified Soc, Co, and Hung as breeds warranting particular attention in conservation planning due to elevated autozygosity. Because each breed was represented by only six individuals, and because no individual-level phenotypic measurements were available, all findings are reported as exploratory population-genomic signals requiring replication in larger cohorts. Overall, we describe genomic differentiation and candidate selection signatures among local Vietnamese pig breeds and provide a foundation for further genomic studies of these breeds.

Van Thanh Nguyen, Ba Văn Nguyễn, D. N. Do · 0 citations
Open access Aug 2026

Genome Wide Structural Variants Provide Insights Into Population Structure and Genetic Divergence in Pacific White Shrimp ( Penaeus vannamei ) Breeding Populations

Structural variants (SVs) are a major yet underused source of adaptive variation in aquaculture. We built a genome‐wide SV atlas for 180 Penaeus vannamei from six commercial breeding populations and discovered 1,159,046 SVs, with uneven chromosomal distributions and multi‐type hotspots. Over 63.53% of SVs overlapped repeats—especially simple sequence repeats, DNA transposons, and LINEs. SV and SNP densities were highly correlated. Across populations, 482 k SVs were shared and 145,623 were singletons; the fraction of deletions increased from shared to singleton classes. BMK and KH harbored more singletons than SIS, RH, and CP, indicating greater divergence. PCA and ADMIXTURE recovered three major clusters and revealed the substructure in RH, mirroring SNP analyses. Selection scans identified 78–193 sweep windows per population encompassing 38–161 candidate genes. These genes were predominantly enriched in population‐specific processes such as chromatin regulation, meiotic recombination, membrane‐associated functions, suggesting that structural variants may contribute to divergence in reproductive, metabolic, and structural pathways across breeding programs. Nevertheless, 10 genes showed parallel signals in over 3 populations; many carry short deletions likely affecting regulatory or coding elements. Together, these results show that genome architecture and domestication jointly shape the shrimp SV landscape; that SVs alone robustly resolve population history; and that a small set of recurrent, deletion‐bearing regulatory genes may underpin convergent improvement. The SV map and candidate loci provide diagnostic markers for germplasm tracing and candidate loci for marker‐assisted or genomic selection in P. vannamei breeding.

Ming-Yang Zhao, Hao Wang, Mingxuan Teng et al. · 0 citations
Open access Sep 2026

Low-pass whole-genome sequencing reveals genomic diversity and ecotype-specific adaptation in indigenous Tigrayan chickens

Indigenous chickens play a critical role in food security and climate resilience in smallholder systems, yet their genomic diversity and adaptive potential remain insufficiently characterised. This study employed low-pass whole-genome sequencing (LP-WGS; 0.2–1.99×) to investigate genomic diversity, population structure, inbreeding and candidate environment-associated genomic variation in 33 chickens from highland, midland, and lowland agroecologies in the Tigray region of northern Ethiopia. After imputation and stringent filtering, 23.4 million high-confidence SNPs were retained, including ~ 17% novel variants, indicating substantial uncharacterised genetic diversity in these populations. SNP density (13.8 ± 8.6 SNPs/kb) was comparable to values reported from high-coverage Ethiopian chicken datasets, demonstrating the suitability of LP-WGS for population genomics in resource-limited settings. Marked differences in genomic diversity were observed among ecotypes: midland chickens showed the highest nucleotide diversity (π = 0.00267), followed by lowland (π = 0.00233), whereas highland chickens showed the lowest diversity (π = 0.00203) and elevated genomic inbreeding (FROH and FHOM ≈ 0.18). Population structure analyses revealed clear genetic separation among ecotypes. PCA (13.91% variation explained) distinguished lowland chickens along PC1 and separated highland from midland along PC2, while ADMIXTURE and FST patterns supported three major ancestral genomic backgrounds. Functional annotation of private missense variants uncovered distinct adaptive signatures reflecting the contrasting agroecological conditions. Highland chickens showed enrichment of candidate genes potentially involved in physiological processes relevant to high-altitude environments, including cold response, angiogenesis, cardiovascular regulation and metabolic homeostasis (eg., PARP1, ACOX2, ITGB3, EDNRB, SOX8, and SOX10). Midland chickens exhibited candidate signals of selection in genes with known roles in innate antiviral immunity, bacterial defence and inflammatory regulation (eg., BAK1, CLSTN1, CYSLTR1, CYSLTR2, CXCR7, GIPR, DSCAM, GDAP1, TLR3, TLR4, TLR7, IFIH1, ADORA1, EPHB1, and TMPRSS2). Lowland chickens displayed candidate variants associated with heat-stress response, DNA damage repair, oxidative balance and cardiovascular support under extreme temperatures (e.g., MLH1, BDKRB1, GPR19, FLT1, CCL18, TGM2, and RAMP3). Overall, the results indicate substantial genomic differentiation among ecotypes and suggest candidate environment-associated genetic divergence across Tigray’s diverse agroecological zones. These populations may represent important reservoirs of adaptive genetic variation for climate-resilient poultry breeding, warranting further functional validation and conservation-oriented management.

Gebreslassie Gebru, G. Belay, Tsadkan Zegeye et al. · 0 citations
Open access Jul 2026

GWAS and selection signature analyses of whole-genome sequencing data identify genes associated with body size in Duolang sheep

Introduction Duolang sheep is a meat-fat dual-purpose breed originating from Xinjiang, China, and the genetic basis of body size variation in this breed remains unclear. Methods Whole-genome resequencing was performed on 224 Duolang sheep with complete records for six body size traits. After variant calling, imputation, and quality control, 19.65 million high-quality SNPs were retained. After linkage disequilibrium pruning, 2.1 million relatively independent SNPs were used for a mixed linear model GWAS in GEMMA. Tajima’s D and integrated haplotype score analyses were performed in a kinship-filtered subset of 60 animals, and SheepGTEx data were used to characterize tissue expression patterns. Results The GWAS identified 24 significant SNPs for height, body length, chest girth, shoulder width, and hip width. Integration of GWAS intervals with selection-signal regions prioritized 31 candidate genes, including CA10, WDR45B, FADS2, LGALS12, and CTNNA2. Multi-tissue expression profiles provided functional context for the prioritized genes. Discussion These findings provide insight into the genetic architecture of body size traits in Duolang sheep and identify candidate loci for future validation and genetic improvement studies.

Keyao Wang, Zhigang Niu, Sen Tang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.