Aug 2026· Nature Genetics· Vol 58, pp. 1941 - 1952· 1 citation· 93 references
Medicine
TL;DR
A family of classification models, scE2G, is introduced that predict enhancer–gene regulatory interactions from single-cell datasets and enable mapping of these interactions across diverse cell types and tissues and will enable accurate mapping of enhancer–gene regulatory interactions across thousands of human cell types.
Abstract
Mapping enhancers and their target genes in specific cell types is crucial for understanding gene regulation and human disease genetics. However, accurately predicting enhancer–gene regulatory interactions from single-cell datasets has been challenging. Here we introduce a family of classification models, scE2G, to predict enhancer–gene regulation. These models use features from single-cell assay for transposase-accessible chromatin with sequencing (ATAC-seq) or multiomic RNA and ATAC-seq data, and are trained on a CRISPR perturbation dataset including >10,000 evaluated element–gene pairs. We benchmark scE2G models against CRISPR perturbations, fine-mapped expression quantitative trait loci and genome-wide association study variant–gene associations and demonstrate state-of-the-art performance at prediction tasks across several cell types and categories of perturbations. We apply scE2G to build maps of enhancer–gene regulatory interactions in heterogeneous tissues and interpret noncoding variants associated with complex traits, nominating regulatory interactions linking INPP4B and IL15 to lymphocyte count. The scE2G models will enable accurate mapping of enhancer–gene regulatory interactions across thousands of human cell types. scE2G is a family of models that predict enhancer–gene regulatory interactions from single-cell datasets and enable mapping of these interactions across diverse cell types and tissues.
Genetic effects on complex traits primarily act by regulating gene expression; however, this process is not well understood1. Studies of genetic effects on gene expression (expression quantitative trait loci (eQTLs)) can inform as to the gene regulatory layer between genetic variants and complex traits2. However, previous studies have not effectively captured cell-type-specific eQTLs, which are likely to be important for complex traits. Here we unbiasedly characterized cell-type-specific eQTLs by applying a variance component model to population-scale single-cell RNA-sequencing (RNA-seq) data. Using peripheral blood mononuclear cells from the OneK1K cohort, we demonstrated that cell-type-specific eQTLs enrich for complex trait heritability, which we did not observe for cell-type-shared eQTLs. We also found that eQTL specificity is associated with genes that have greater selective constraint, enhancer complexity and gene network connectivity, three features enriched in complex traits relative to known eQTLs3,4. Transcriptome-wide, trans eQTLs were mostly cell-type-specific (60% specific) whereas cis eQTLs were mostly shared (30% specific). We used a second single-cell RNA-seq dataset to replicate our findings and demonstrate that cell-type-shared and cell-type-specific eQTLs are consistent across ancestries. Our results establish eQTL cell-type specificity as a key feature of gene regulation and partly explain why known eQTLs are depleted in gene regulatory effects on complex traits.
Minhui Chen, Xin-Pei Wang, Lena Krockenberger et al.· Nature· 0 citations
Understanding gene regulation at single-cell resolution is crucial for unraveling development, disease, and cellular identity. We introduce single-cell regulatory graph attention network (scReGAT), a deep learning framework that integrates prior knowledge of cis-regulatory element (cRE)-gene and transcription factor-gene interactions to reconstruct cell-specific regulatory networks. Central to scReGAT is a knowledge-guided regulatory graph (kRG), which combines experimentally validated regulatory interactions with cell-resolved chromatin accessibility profiles. These graphs serve as the foundation for training a Graph Attention Network (GAT) to predict gene expression and quantify the contribution of specific regulatory interactions using an interpretable regulatory score for each edge. In benchmarking across five single-cell multi-omics datasets, scReGAT successfully recapitulates known cell-type-specific cRE-gene interactions. In both neuroblastoma and osteogenic differentiation systems, it uncovers dynamic regulatory rewiring that predicts transcriptional transitions. Furthermore, by integrating genome-wide association studies loci from Alzheimer's disease, multiple sclerosis, and schizophrenia, scReGAT identifies disease-associated cell types and uncovers candidate regulatory mechanisms underlying complex trait associations. These results position scReGAT as a robust and generalizable framework for decoding long-range gene regulation at single-cell resolution. The source code of scReGAT can be accessed at https://github.com/TianLab-Bioinfo/scReGAT/ and https://ngdc.cncb.ac.cn/biocode/tool/BT008081.
Gene regulatory networks encode the fundamental logic of cellular functions, but systematic network mapping remains challenging, especially in cell states relevant to human biology and disease. Here, we perturbed all expressed genes across 22 million primary human CD4+ T cells from four donors and developed a probe-based perturb-seq platform to measure the transcriptome effects in cells at rest and after stimulation. These data allowed us to map genes regulating immune pathways, including previously uncharacterized regulators of cytokine production. Importantly, active regulators and the gene programs they control changed dramatically across stimulation conditions. Perturbation signatures enabled us to model T cell states observed in population-scale transcriptomic atlases, nominating regulators of T cell polarization and of age-related phenotypes. Finally, we leveraged perturb-seq to implicate context-specific gene regulatory pathways in autoimmune disease risk. Our study provides a foundational resource and new approaches to decode T cell function and human immune traits.
Ronghui Zhu, E. Dann, Jun Yan et al.· Cell· 0 citations
MOTIVATION
Predicting and deciphering the regulatory logic of enhancers remains a significant challenge due to their complex sequence features and the absence of consistent genetic or epigenetic signatures that distinguish them from other genomic regions. Existing machine learning methods capture nucleotide composition but often fail to model sequence context effectively.
RESULTS
We present DNABERT-Enhancer, a novel enhancer prediction method, by applying DNABERT pre-trained language model on the human genome. Using ENCODE registry of candidate cis-regulatory elements (cCREs), we curated a benchmark dataset, consisting of 21,926 enhancers of 201 bp length and 46,159 enhancers of 350 bp length, as positive instances. The best fine-tuned model achieved 88.05% accuracy and a Matthews correlation coefficient of 76% on an independent dataset. Genome-wide application identified 1,684,595 enhancer regions covering 26.65% of the human genome. By performing integrative analyses with DNABERT-based transcription factor models, we identify 2,681 statistically significant loss-of-function and 1,917 gain-of-function enhancer variants, which respectively alter the function of 1,623 and 1,247 ENCODE-cCRE enhancers. Similarly, we identify 4,057 candidate de novo enhancers, created by 5,464 gain-of-function variants. These genome-wide enhancer annotations and candidate genetic variants predicted by DNABERT-Enhancer provide valuable resources for genome interpretation in functional and clinical genomics studies.
AVAILABILITY
DNABERT-Enhancer is freely available at https://github.com/DavuluriLab/DNABERT-Enhancer; Trained model predictions can be explored interactively via the web application at https://dnabert-enhancer-datarepo.streamlit.app/. The fine-tuned models are archived and citable through Zenodo (https://doi.org/10.5281/zenodo.19157566).
SUPPLEMENTARY INFORMATION
Supplementary data are available at Bioinformatics online.
R. Sathian, P. Dutta, F. Ay et al.· Bioinformatics· 0 citations
Sequence-to-function models learn regulatory features from genomic sequence, but they remain limited in their ability to predict gene-expression differences among individuals. Cell-type-specific regulatory effects may be obscured in bulk RNA sequencing, whereas paired genotype and single-cell expression cohorts remain small. We evaluated whether deconvolution of bulk RNA-seq could provide scalable cell-type-specific targets for personal-genome expression prediction. GTEx v8 bulk RNA-seq from six tissues was deconvolved with BayesPrism using single-nucleus reference profiles, producing targets across 83 tissue–cell-type contexts. Deconvolved expression agreed with matched pseudobulked GTEx single-nucleus RNA-seq, with median donor-level Pearson correlations across genes ranging from 0.53 to 0.73 by tissue. We compared genotype-feature models, regressors trained on frozen Enformer representations, and fine-tuned Enformer and Borzoi models. Across random and nonlinear-enriched gene sets, sequence-derived approaches generally outperformed genotype-feature baselines, while frozen Enformer features were competitive with end-to-end fine-tuning. For the random gene set, Fisher-averaged Pearson correlations were 0.122–0.142 for sequence-derived approaches and 0.081–0.086 for genotype-feature baselines in a coverage-aware sensitivity analysis. Model performance was positively associated with deconvolution–pseudobulk agreement for sequence-derived models (r = 0.35–0.43 across tissue–cell-type contexts), suggesting that target reliability may constrain downstream prediction. Context-specific Enformer fine-tuning did not materially out-perform a shared, combined-context strategy. These results support deconvolution as a feasible approach for generating cell-type-resolved training targets, while showing that target quality and limited cohort size remain important constraints. Frozen pretrained representations provide a computationally efficient and competitive baseline for personal sequence-to-expression modeling.
Motivation Single-cell expression quantitative trait locus (eQTL) studies can resolve cell-type-specific genetic effects, but conventional gene-by-gene analyses do not directly capture coordinated genetic regulation of neighboring genes. Principal-component QTL (pcQTL) mapping can summarize such multi-gene effects, but existing approaches were developed for bulk expression and are not designed for sparse single-cell counts. Results: We developed sc-pcQTL, a framework that applies two-component hurdle modeling and sliding-window clustering to identify local co-expression clusters, summarizes each cluster using principal components, and maps cis-pcQTLs. In simulations, the individual hurdle components controlled type I error, while the component-union screening rule was substantially more powerful than donor-level pseudobulk correlation tests. Applied to 1.24 million peripheral blood mononuclear cells from 982 OneK1K donors across 10 cell types, sc-pcQTL identified 2,485 local co-expression clusters and conducted QTL mapping for 4,353 cluster-PC phenotypes at single-cell resolution, of which 2,040 had at least one significant cis-pcQTL association. Fine-mapping and colocalization with genome-wide association study loci across 1,163 phenotypes in the FinnGen study identified 394 colocalized QTL-GWAS signal groups. Each group comprised fine-mapped QTL and GWAS signals connected through one or more colocalization links within the same cell type and local gene cluster. Of these groups, 46 were pcQTL-specific and contained no colocalized single-gene eQTL from a constituent gene. Locus-level analyses further revealed cell-type-specific multi-gene regulatory effects. Thus, sc-pcQTL complements conventional single-gene eQTL analysis by identifying trait-relevant regulatory signals shared across neighboring genes. Availability and implementation: The sc-pcQTL software is openly available at https://github.com/ZhouLabGenetics/sc-pcQTL; analysis and figure-generation scripts are available at https://github.com/ZhouLabGenetics/sc-pcQTL_code; and the summary result tables are publicly available on Zenodo (DOI: https://doi.org/10.5281/zenodo.21222687). Contact wzhou@broadinstitute.org Supplementary material Supplementary material accompanies this preprint.
Jun-Kai Zhang, Yi Huang, M. Claussnitzer et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.