An encyclopedia of enhancer–gene regulatory interactions in the human genome is built, revealing global properties of enhancer networks, identifying differences in regulatory complexity across genes, and improving analyses linking noncoding variants to target genes and cell types for common, complex diseases.
Abstract
Identifying transcriptional enhancers and their target genes is essential for understanding gene regulation and the effect of human genetic variation on disease1, 2, 3, 4, 5–6. Here we create and evaluate a resource of more than 92 million enhancer–gene regulatory interactions across 1,458 biosamples covering 369 cell types and tissues, by integrating predictive models, chromatin states, three-dimensional contacts and large-scale genetic perturbations generated by the ENCODE Consortium7. We first create a systematic benchmarking pipeline to compare predictive models, assembling a dataset of 10,356 element–gene pairs measured in CRISPR perturbation experiments, more than 30,000 fine-mapped expression quantitative trait loci and 569 fine-mapped genome-wide association study (GWAS) variants linked to a probable causal gene. Using this framework, we develop ENCODE-rE2G, a predictive model achieving state-of-the-art performance across several prediction tasks, demonstrating that iterative perturbations and supervised machine learning can build increasingly accurate predictive models of enhancer regulation. Using ENCODE-rE2G, we build an encyclopedia of enhancer–gene regulatory interactions in the human genome, revealing global properties of enhancer networks, identifying differences in regulatory complexity across genes and improving analyses linking noncoding variants to target genes and cell types for common complex diseases. By interpreting the model, we find that beyond enhancer activity and three-dimensional enhancer–promoter contacts, additional features that guide enhancer–promoter communication include promoter class and enhancer–enhancer synergy. These genome-wide maps of enhancer–gene regulatory interactions, benchmarking software, predictive models and insights about enhancer function provide a valuable resource for future studies of gene regulation and human genetics. An encyclopedia of more than 92 million enhancer–gene regulatory interactions created as part of the ENCODE4 project provides a valuable resource for future studies of gene regulation and human genetics.
A family of classification models, scE2G, is introduced that predict enhancer–gene regulatory interactions from single-cell datasets and enable mapping of these interactions across diverse cell types and tissues and will enable accurate mapping of enhancer–gene regulatory interactions across thousands of human cell types.
Maya U. Sheth, Wei-Lin Qiu, X. Ma et al.· Nature Genetics· 1 citation
Understanding gene regulation at single-cell resolution is crucial for unraveling development, disease, and cellular identity. We introduce single-cell regulatory graph attention network (scReGAT), a deep learning framework that integrates prior knowledge of cis-regulatory element (cRE)-gene and transcription factor-gene interactions to reconstruct cell-specific regulatory networks. Central to scReGAT is a knowledge-guided regulatory graph (kRG), which combines experimentally validated regulatory interactions with cell-resolved chromatin accessibility profiles. These graphs serve as the foundation for training a Graph Attention Network (GAT) to predict gene expression and quantify the contribution of specific regulatory interactions using an interpretable regulatory score for each edge. In benchmarking across five single-cell multi-omics datasets, scReGAT successfully recapitulates known cell-type-specific cRE-gene interactions. In both neuroblastoma and osteogenic differentiation systems, it uncovers dynamic regulatory rewiring that predicts transcriptional transitions. Furthermore, by integrating genome-wide association studies loci from Alzheimer's disease, multiple sclerosis, and schizophrenia, scReGAT identifies disease-associated cell types and uncovers candidate regulatory mechanisms underlying complex trait associations. These results position scReGAT as a robust and generalizable framework for decoding long-range gene regulation at single-cell resolution. The source code of scReGAT can be accessed at https://github.com/TianLab-Bioinfo/scReGAT/ and https://ngdc.cncb.ac.cn/biocode/tool/BT008081.
Regulatory genomics faces a depth–breadth gap: deep multi-omics provides regulatory detail but is difficult to scale, whereas broad expression datasets often lack the regulatory structure needed for mechanistic Gene Regulatory Network (GRN) analysis. We developed Regulatory Elements Guided Analysis (REGA), an interpretable framework that uses reference Regulatory Element (RE) catalogs to infer transcription factor (TF)–RE–gene programs from gene expression data. Across ChIP-seq, knockdown, Hi-C, cis- and trans-eQTL benchmarks, REGA prioritized functional REs, improved RE–gene and TF–gene inference over existing baselines, including methods using more data, and recovered coherent regulatory modules. In PsychENCODE snRNA-seq, REGA identified disease-associated modules and TF activities, linked regulatory dysregulation to genetic risk, and detected cross-cell-type neuronal–glial programs. In spatial transcriptomics, REGA linked cell-intrinsic regulatory programs with intercellular ligand–receptor communication; in Perturb-seq, it mapped perturbation responses to trait-associated regulatory architectures. REGA enables scalable, interpretable GRN analysis across expression datasets.
This work provides a foundation for applications that link epigenome variation to gene expression in human cells, by benchmarking methods on a per-gene basis, illustrating their use in a disease context and making trained models available to the community.
Fatemeh Behjati Ardakani, Shamim Ashrafiyan, Laura Rumpf et al.· Genome Biology· 0 citations
This review compares convolutional, Transformer-based and graph architectures used to represent local sequence features, chromatin state and three-dimensional genome organisation to their applications to transcription-factor binding, chromatin accessibility, gene expression, non-coding variant prioritisation and regulatory-sequence design.
Deciphering the regulatory consequences of sequence divergence across human evolution is essential to understanding the molecular basis of human-specific traits and disease. Although millions of derived alleles distinguish humans from great apes, only a small fraction are likely to influence human-specific traits. Previous studies have focused on regions of elevated sequence divergence, assuming that rapid evolution reflects functional adaptation, yet individual high-impact regulatory mutations evade such scans. Here, we apply sequence-to-function deep learning to predict chromatin accessibility across modern human, archaic hominin, and great ape personalized genomes, identifying lineage-specific cis-regulatory elements (linCREs) across diverse cellular contexts. Compared to conserved elements, linCREs are shorter, less pleiotropic, less conserved, and enriched in neurodevelopmental pathways. Many linCREs occur in regions with limited sequence divergence that acceleration-based approaches would overlook. We validate lineage-specific enhancer activity through luciferase reporter assays and demonstrate that a single motif-generating derived allele nominated by model interpretability tools drives a hominin-specific neurodevelopmental enhancer.
Riley J. Mangan, Nikitha Thoduguli, Dimitar Ivanov et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.