Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Hidden genetic diversity in 320 nearly-complete East Asian genome assemblies

East Asian populations, representing over 20% of the global population, remain critically underrepresented in human genomic studies, limiting our understanding of population-stratified genetic variation and its implications for health and disease. Here we present the first phase of the Asian Pan-Genome project (APG), comprising 320 nearly complete, fully phased haploid genome assemblies from 160 East Asian individuals. These assemblies achieve unprecedented quality, with an average contig N50 of 144.3 megabase pairs and an average quality value of 64.5. Leveraging these superior assemblies, we reveal previously uncharacterized diversity in human repeatome, including population-stratified patterns in centromere satellites and rDNA arrays. Compared to existing global human genome assemblies, the newly generated genomes supplement 152 million base pairs of novel sequences, 355 gene gains, 18,300 structural variation loci and 26 large euchromatic inversions missing from current human pangenomes. We perform population stratification analyses of structural variations, and further resolve the structural haplotypes of complex genomic regions such as Major Histocompatibility Complex and Survival Motor Neuron loci across global pangenomes, exemplifying tandem-duplicate and inversion-rich complex locus architectures in the human genome, respectively. This resource provides a critical foundation for human genetic studies, especially for East Asian populations, promoting more accurate variant discovery, reducing bias, and ultimately advancing the equity and efficacy of genomic medicine.

Dongya Wu, Chentao Yang, Quanyu Chen et al. · 2 citations
Open access Aug 2026

Dynamic sample augmentation discovers gene biomarker for IgA nephropathy

Bulk RNA-seq data suffers from the issues of “high dimensionality and small sample size,” which limits its application in disease research. This paper proposes a dynamic data augmentation method based on Layer-wise Relevance Propagation (LRP) aimed at improving classification performance and biological interpretability under small-sample conditions. The method utilizes the LRP algorithm to calculate the contribution weight of each gene to the classification result and uses this weight to guide sample generation. By systematically amplifying biologically meaningful signals, it constructs semantically reliable augmented samples, avoiding the semantic distortion caused by traditional random perturbations. Simultaneously, a dynamic augmentation mechanism is introduced that tightly couples sample generation with model training, providing difficult-to-classify samples with multiple iterative optimization opportunities and forming a virtuous cycle where classification performance and augmentation quality improve synergistically. On this basis, population-level gene biomarkers are identified from the trained model. Innovatively, an open-environment enrichment analysis method is proposed—that is, instead of being limited to a few feature genes of a single subtype, the union of feature genes from all disease subtypes is taken for enrichment analysis, revealing shared biological pathways from a systems-level perspective and providing a more comprehensive interpretation for subtype-specific mechanism research. Experimental results show that this method effectively improves classification accuracy, and through this open-environment enrichment approach, six hub genes were identified as gene markers for IgAN.

Fang Zheng, Juanjuan Zhao, Baoping Jia et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.