Skip to content
Open access

Sequence-to-function deep learning decodes human cis-regulatory evolution

Aug 2026 · bioRxiv · 0 citations
Biology

Abstract

Deciphering the regulatory consequences of sequence divergence across human evolution is essential to understanding the molecular basis of human-specific traits and disease. Although millions of derived alleles distinguish humans from great apes, only a small fraction are likely to influence human-specific traits. Previous studies have focused on regions of elevated sequence divergence, assuming that rapid evolution reflects functional adaptation, yet individual high-impact regulatory mutations evade such scans. Here, we apply sequence-to-function deep learning to predict chromatin accessibility across modern human, archaic hominin, and great ape personalized genomes, identifying lineage-specific cis-regulatory elements (linCREs) across diverse cellular contexts. Compared to conserved elements, linCREs are shorter, less pleiotropic, less conserved, and enriched in neurodevelopmental pathways. Many linCREs occur in regions with limited sequence divergence that acceleration-based approaches would overlook. We validate lineage-specific enhancer activity through luciferase reporter assays and demonstrate that a single motif-generating derived allele nominated by model interpretability tools drives a hominin-specific neurodevelopmental enhancer.

Read PDF

Similar papers

Review Open access Aug 2026

Deep Learning for Deciphering the Plant Cis-Regulatory Code

This review compares convolutional, Transformer-based and graph architectures used to represent local sequence features, chromatin state and three-dimensional genome organisation to their applications to transcription-factor binding, chromatin accessibility, gene expression, non-coding variant prioritisation and regulatory-sequence design.

Zhi-Meng Zhao, Si-Xuan Huang, Shi-Long Zhang et al. · 0 citations
Open access Sep 2026

Genomic language model for predicting enhancers and their allele-specific activity in the human genome.

MOTIVATION Predicting and deciphering the regulatory logic of enhancers remains a significant challenge due to their complex sequence features and the absence of consistent genetic or epigenetic signatures that distinguish them from other genomic regions. Existing machine learning methods capture nucleotide composition but often fail to model sequence context effectively. RESULTS We present DNABERT-Enhancer, a novel enhancer prediction method, by applying DNABERT pre-trained language model on the human genome. Using ENCODE registry of candidate cis-regulatory elements (cCREs), we curated a benchmark dataset, consisting of 21,926 enhancers of 201 bp length and 46,159 enhancers of 350 bp length, as positive instances. The best fine-tuned model achieved 88.05% accuracy and a Matthews correlation coefficient of 76% on an independent dataset. Genome-wide application identified 1,684,595 enhancer regions covering 26.65% of the human genome. By performing integrative analyses with DNABERT-based transcription factor models, we identify 2,681 statistically significant loss-of-function and 1,917 gain-of-function enhancer variants, which respectively alter the function of 1,623 and 1,247 ENCODE-cCRE enhancers. Similarly, we identify 4,057 candidate de novo enhancers, created by 5,464 gain-of-function variants. These genome-wide enhancer annotations and candidate genetic variants predicted by DNABERT-Enhancer provide valuable resources for genome interpretation in functional and clinical genomics studies. AVAILABILITY DNABERT-Enhancer is freely available at https://github.com/DavuluriLab/DNABERT-Enhancer; Trained model predictions can be explored interactively via the web application at https://dnabert-enhancer-datarepo.streamlit.app/. The fine-tuned models are archived and citable through Zenodo (https://doi.org/10.5281/zenodo.19157566). SUPPLEMENTARY INFORMATION Supplementary data are available at Bioinformatics online.

R. Sathian, P. Dutta, F. Ay et al. · 0 citations
Review Jul 2026

Toward generalizable and interpretable AI in regulatory genomics

It is suggested that progress requires reframing seq2func models as continually refined systems, in which targeted perturbation experiments, systematic evaluation and iterative model updates are tightly coupled through artificial intelligence-experiment feedback loops, enabling self-improving models that progressively deepen mechanistic understanding and more reliably support biological discovery.

Masayuki Nagai, A. E. Murphy, Kaeli Rizzo et al. · 2 citations
Open access Sep 2026

Deep Learning-Guided Identification and In Vivo Validation of Compact Cis-Regulatory Elements for the Zebrafish Habenula

Background/Objectives: Precise genetic access to the zebrafish habenula remains limited by a scarcity of compact, sequence-defined cis-regulatory elements (CREs). Here, we integrated developmental expression mapping, deep-learning predictions on long-range sequences, and in vivo reporter assays to identify compact regulatory sequences driving habenular expression. Methods: Using a transgenic zebrafish line enriched for habenular reporter expression, we isolated GFP-positive cells from larval brains and profiled their transcriptomes via microarray. A subset of candidate genes enriched in this dataset was validated using whole-mount in situ hybridization across two developmental stages. This analysis identified genes with highly reproducible habenular expression, leading to the selection of the gng8 and ano2 loci for subsequent CRE characterization. We developed ZEN-former (Zebrafish EN-former), an Enformer-based sequence-to-function model trained on neuronal subclass chromatin accessibility profiles from the adult mouse brain. Results: The model demonstrated strong correlation between predicted and experimentally measured signals across held-out genomic regions. To prioritize regulatory candidates, we integrated ZEN-former predictions with available zebrafish ATAC-seq data, RepeatMasker annotations, and gene models, identifying two ~600 bp intervals at each gene locus. These selected intervals were combined to generate ~1.2 kb reporter constructs for gng8 and ano2 loci, which were then evaluated using Tol2 transposon mediated transgenesis assays in zebrafish. In transiently injected larvae, both constructs successfully drove reporter expression in the habenular region. Furthermore, the resulting stable transgenic lines displayed highly specific and reproducible habenular expression. Quantitative confocal analysis showed mean habenular labeling completeness values of 84.5% and 96.2% for the gng8- and ano2-derived lines, respectively. Conclusions: Together, these findings provide a proof of concept that sequence features learned from mammalian chromatin accessibility datasets can effectively guide the prioritization of functional regulatory elements across species in zebrafish. The compact regulatory constructs and stable transgenic lines generated here offer robust genetic tools for investigating habenular circuitry.

Ze-Ran Li, Shan-Shan Liu, Quan Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.