Skip to content
Open access

Chromosome-scale genome assembly and annotation of the Vietnamese indica rice cultivar Khang Dan 18

Aug 2026 · bioRxiv · 0 citations · 45 references
Biology

TL;DR

A chromosome-scale genome assembly of KD18 generated using Oxford Nanopore long-read and Illumina short-read sequencing is reported, providing a cultivar-specific genomic resource for research involving KD18 and KD18-derived materials.

Abstract

Khang Dan 18 (KD18) is an Oryza sativa L. subsp. indica rice cultivar widely cultivated in northern Vietnam and used as an experimental and breeding background in Vietnamese rice research. Although KD18 has previously been represented in low-depth population resequencing datasets, a contiguous and annotated cultivar-specific genome has not been available. Here, we report a chromosome-scale genome assembly of KD18 generated using Oxford Nanopore long-read and Illumina short-read sequencing. The 395.3-Mb assembly comprises 12 chromosome-scale pseudomolecules containing approximately 95% of the assembled sequence and 99.6% of the predicted protein-coding genes. The assembly showed 97.2% BUSCO completeness, an average Merqury quality value of 46 and a long terminal repeat assembly index of 13.21. A total of 56,546 protein-coding genes representing 71,237 transcripts were predicted, with 99% BUSCO and 98.68% OMArk completeness. These statistics are similar to those of other high-quality genome assemblies that were recently published for different Asian rice cultivars, therefore providing a cultivar-specific genomic resource for research involving KD18 and KD18-derived materials.

Read PDF

Similar papers

Open access Aug 2026

Reference-guided comparative genomics of seven Indonesian rice cultivars identifies conserved gene space and trait-associated sequence candidates

Indonesian rice cultivars represent valuable genetic resources, yet many remain poorly characterized at the genomic level. Here, we generated 95.40 Gb of PacBio HiFi sequence data from seven Indonesian rice cultivars and constructed cultivar-specific consensus genomes using the telomere-to-telomere Nipponbare reference AGIS1.0. Sequencing coverage ranged from 27.92× to 41.58×, and the resulting consensus genomes spanned 387.93-390.54 Mb, with BUSCO completeness of approximately 98.3-98.5%. OrthoFinder assigned 99.1% of predicted proteins to 40,737 orthogroups, including 27,514 core orthogroups represented across all seven cultivars, indicating a highly conserved predicted gene space within the reference-guided framework. Targeted analysis recovered 278 of 280 cultivar-by-locus combinations representing 40 genes or gene family entries associated with grain pigmentation, nitrogen and amino-acid metabolism, and starch properties. Comparative predicted protein analysis prioritized ANS1, SBE2b, SSIIa/ALK, Wx/GBSSI, OsAAP6/qPC1, and SSI as candidates for further investigation. Among 269 completed AGIS1.0-anchored promoter comparisons, 159 passed quality-control criteria, whereas 110 were flagged for gene-model, boundary, synteny, or structural concerns. Notably, these flagged comparisons accounted for more than 90% of the alignment-derived sequence variation, emphasizing the importance of rigorous quality control when interpreting apparent promoter divergence. Collectively, these reference-guided genomic resources provide a standardized framework for investigating sequence variation in Indonesian rice germplasm and prioritize testable coding and regulatory candidates for functional validation and future genomics-assisted crop improvement.

Y. Purwestri, Adhityo Wicaksono, Siti Nurbaiti et al. · 0 citations
Open access Aug 2026

Chromosome-level genome assembly of tea cultivar Fuding Dahao

Tea ( Camellia sinensis ) is a globally important economic crop. Among elite cultivars, ‘Fuding Dahao’ is particularly prized for its superior agronomic traits and its central role in premium white tea production. However, the lack of a high-quality chromosome-level genome for this regionally adapted cultivar has hindered molecular breeding efforts. Here, we present the first chromosome-level reference genome of ‘Fuding Dahao’ assembled using PacBio HiFi long-read sequencing and Hi-C chromatin interaction mapping. The final 3.30 Gb assembly is highly contiguous, with 90% of the sequences anchored to 15 pseudochromosomes. A total of 54,345 protein-coding genes were predicted, representing a substantial improvement in both assembly contiguity and annotation completeness compared with previously published tea genomes. This high quality genome provides a critical resource for dissecting the genetic basis of white tea quality traits and accelerating molecular breeding programs. Our results fill a major gap in tea genomics and lay a solid foundation for the development of superior, locally adapted tea cultivars.

Yang Chen, Lizhong Wang, Deng-Feng Shen et al. · 0 citations
Open access Aug 2026

A subgenome-resolved and chromosome-scale reference genome assembly of allotetraploid wheat wild relative Aegilops peregrina

Aegilops peregrina is a wild allotetraploid wheat wild relative and an important source of genetic diversity for stress tolerance and agronomic traits. Here, we report a subgenome-resolved, chromosome-scale reference genome assembly of a drought tolerant and stem rust resistant Ae. peregrina accession PI 604178 generated using PacBio HiFi and Hi-C sequencing. The 10.13 Gb assembly contains 98.81% of sequence anchored to 14 pseudomolecules representing the seven Sᵖ and seven Uᵖ chromosomes, with contig and scaffold N50 values of 25.84 and 746.48 Mb, respectively. The assembly achieved a consensus quality value of 74.61, 97.83% k-mers completeness, and 99.9% BUSCO completeness. LTR Assembly Index values of 20.43 and 18.79 for the Sᵖ and Uᵖ subgenomes, respectively, further supported high continuity across repeat-rich regions. Repetitive elements comprise 85.93% of chromosome-anchored assembly. We annotated 59,910 high-confidence protein-coding genes, with comparable gene representation across the two subgenomes. This reference genome provides a high-quality genomic framework for comparative analyses, characterization of important loci regulating agronomic and resilience related traits, and sequence-guided exploitation of Ae. peregrina allelic diversity for wheat improvement.

Jatinder Singh, Santosh Gudi, P. Maughan et al. · 0 citations
Open access Aug 2026

Chromosome‐scale assembly of wheat cultivar Sumai 3, a major germplasm source for Fusarium head blight resistance

Abstract Fusarium head blight (FHB) is a devastating disease that severely impacts global wheat (Triticum aestivum L.) production. Sumai 3, a wheat cultivar widely used in breeding programs for its strong FHB resistance, has not been fully resolved at the chromosome level. Here, we present a high‐quality chromosome‐scale assembly of Sumai 3 using PacBio HiFi reads and chromosome conformation capture sequencing. The 14.6 Gb assembly consists of 832 contigs, with the longest contig being 245.2 Mb and a contig N50 of 41.90 Mb, which were scaffolded into 21 pseudomolecules. De novo annotation identified 104,620 high‐confidence protein‐coding genes and found 92.67% of the genome to consist of repetitive sequences. Synteny analysis showed strong collinearity between Sumai 3 and the wheat reference sequence Chinese Spring IWGSC RefSeq v2.1 (CS). Structural variant analysis identified chromosome 2A with the highest number of deletions (2773) and insertions (2645), while chromosome 3B had the most inversions (359). Duplications were most frequent on 2A (337), and contractions on 5B (122). The gene content of the major resistance quantitative trait loci on 3B, Fhb1, largely validates previous annotations for CS, although we discovered two new genes at approximately 12.4 Mb, including an additional copy of a terpene synthase, further suggesting homology with CS 3D over 3B. Differential expression analysis highlighted up‐regulation of three genes coding pore‐forming toxin‐like protein, sina superfamily protein, and plastid‐lipid‐associated proteins potentially involved in FHB resistance. This assembly provides critical insights into FHB resistance and offers a valuable genomic resource for wheat breeding programs.

Rubylyn D. Mijan, Bikash Poudel, Sittal Thapa et al. · 0 citations
Open access Jul 2026

A highly contiguous genome assembly of Cyclamen persicum to accelerate functional genomics and breeding

Cyclamen is an economically important ornamental plant widely cultivated for its diverse floral characteristics and adaptation to cool climates. Despite its horticultural significance, genomic resources for this species remain limited, hindering molecular studies and genomics-assisted breeding. Here, we report the first highly contiguous nuclear genome assembly of C. persicum generated using high-fidelity long-read sequencing. The assembled genome spans 1.48 Gb, consisting of 126 contigs with an N50 length of 52.3 Mb. Telomeric repeat analysis identified eight contigs containing telomeric sequences at both ends, suggesting the presence of near-complete chromosome assemblies. Genome completeness assessment using BUSCO indicated 98.1% completeness. Repetitive sequences occupied 82.9% of the assembly, with long terminal repeat retrotransposons accounting for 42.1% of the genome. A total of 40,223 protein-coding genes were predicted, with a complete BUSCO score of 95.7%. Comparative orthogroup analysis with five representative eudicot species identified 430 orthogroups specific to C. persicum and 363 orthogroups shared exclusively between C. persicum and Primula kwangtungensis, indicating the presence of both lineage-specific and Primulaceae-conserved gene families. These findings provide critical insights into gene family evolution within Primulaceae and establish an essential comparative framework for future genomic studies. The genome resource presented here provides an invaluable foundation for investigating genome evolution, gene function, and trait-associated loci in cyclamen, effectively facilitating molecular breeding and genetic improvement in this ornamental species.

K. Shirasawa, Y. Akita, Y. Mizunoe et al. · 0 citations
Dataset Open access Aug 2026

Whole-Genome Variants Resource of 144 Oryza rufipogon Accessions

The findings offer a comprehensive genomic variation resource for O. rufipogon, supporting future genetic research and rice breeding efforts, and confirming the data’s reliability.

Li-Hao Zheng, Mohamad Iqbal Hakim Mohd Azhan, S. Ramlee et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.