Skip to content

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Hidden genetic diversity in 320 nearly-complete East Asian genome assemblies

East Asian populations, representing over 20% of the global population, remain critically underrepresented in human genomic studies, limiting our understanding of population-stratified genetic variation and its implications for health and disease. Here we present the first phase of the Asian Pan-Genome project (APG), comprising 320 nearly complete, fully phased haploid genome assemblies from 160 East Asian individuals. These assemblies achieve unprecedented quality, with an average contig N50 of 144.3 megabase pairs and an average quality value of 64.5. Leveraging these superior assemblies, we reveal previously uncharacterized diversity in human repeatome, including population-stratified patterns in centromere satellites and rDNA arrays. Compared to existing global human genome assemblies, the newly generated genomes supplement 152 million base pairs of novel sequences, 355 gene gains, 18,300 structural variation loci and 26 large euchromatic inversions missing from current human pangenomes. We perform population stratification analyses of structural variations, and further resolve the structural haplotypes of complex genomic regions such as Major Histocompatibility Complex and Survival Motor Neuron loci across global pangenomes, exemplifying tandem-duplicate and inversion-rich complex locus architectures in the human genome, respectively. This resource provides a critical foundation for human genetic studies, especially for East Asian populations, promoting more accurate variant discovery, reducing bias, and ultimately advancing the equity and efficacy of genomic medicine.

Dongya Wu, Chentao Yang, Quanyu Chen et al. · 2 citations
Open access Aug 2026

DeepGeSeq: deep learning library for genomic sequence modeling and analysis

Abstract Motivation Deep learning methods have demonstrated significant potential in genomics, enabling broad applications such as sequence activity prediction, regulatory rule identification, and variant effect quantification. However, their widespread adoption is often hindered by the steep computational learning curve required for model construction, training, and downstream biological interpretation. Here, we introduce DeepGeSeq, a user-friendly Deep-learning library tailored for Genomic Sequence modeling and analysis. Results By integrating state-of-the-art architectural modules, DeepGeSeq streamlines the entire deep learning workflow, requiring minimal user input via a simple configuration file and an intuitive agentic skill. We comprehensively validate the efficacy of DeepGeSeq through diverse case studies, encompassing pipeline verification using synthetic datasets, the reproduction and application of established models, and model fine-tuning coupled with biological interpretation on user-defined data. Furthermore, we demonstrate DeepGeSeq’s versatility in domain-specific applications, including single-cell ATAC-seq modeling for cell-type clustering, and MPRA data modeling coupled with in silico saturation mutagenesis to dissect cis-regulatory elements. Ultimately, DeepGeSeq bridges the gap between computational complexity and biological discovery, providing an accessible resource that facilitates the development and broad application of deep learning methods in genomics research. Availability and implementation https://github.com/JiaqiLi1024/DeepGeSeq.

Jiaqi Li · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.