The relationship between DNA sequence and the epigenome relative to RT is dissect and it is suggested that primary DNA sequence instructions exert a powerful baseline control over constitutive and cell-type-specific RT, fine-tuned by regulatory cues beyond the DNA sequence.
DNA replication is a biological process in which a single DNA molecule is duplicated, initiating from multiple genomic sites known as replication origins. Identifying replication origins and analyzing their underlying base sequence composition is crucial for understanding the mechanisms of DNA replication. Although there are various machine learning and deep learning approaches for origin prediction, many rely on labor intensive feature engineering or lack interpretability. We fine-tune two genome-based pretrained language models, DNABERT and DNABERT-2, to predict replication origins in budding yeast and unravel the DNA base composition behind them. The key contribution of this study is a systematic framework for analyzing genomic language models for replication origin prediction, combining controlled dataset design with model-specific explainability pipelines to examine how different tokenization strategies influence learned sequence features and whether such approaches can highlight biologically meaningful signals. We evaluate both models on the designed datasets to ensure robustness and support explainability. DNABERT demonstrates consistent performance, achieving an average accuracy of 0.72 for more challenging and 0.83 for the easier dataset. In comparison, DNABERT-2 achieved comparable scores of 0.72 and 0.81 on the same datasets. Our attention-based motif discovery pipeline enhances the interpretability of DNABERT, by identifying motifs from high-attention fragments that closely match known sequence patterns of replication origins. Perturbation-based explanation methods, including Shapley additive explanations, were applied to interpret DNABERT-2’s learning mechanism. This analysis identified tokens with high attribution scores aligned with biologically relevant sequence composition. Our study demonstrates that both models identify replication origin sequences, albeit through different learning strategies. Tokenization appears to influence model learning and attention behavior in these models. The overlapping k-mer tokenization used in DNABERT yields more interpretable attention maps compared to the byte pair encoding tokenization employed in DNABERT-2. We show that despite sharing the same BERT-style architecture, DNABERT captures relevant short-range patterns and some sequence dependencies beyond just local context, as reflected in its attention maps. In contrast, DNABERT-2’s alternative tokenization strategy biases its learning toward relevant short-range patterns by optimizing token weighting.
Zohreh Piroozeh, I. Akerman, Olga V. Kalinina et al.· BMC Bioinformatics· 0 citations
Transcription factor (TF) occupancy in vivo depends not only on the underlying DNA sequence but also on the local epigenetic environment, which varies across cell types and strongly influences whether sequence-encoded binding potential becomes functional. Here we present EpiBinder, a multimodal deep-learning framework for cell-type-specific prediction of TF binding that jointly models DNA sequence with base-resolution epigenetic information, including cytosine methylation from whole-genome bisulfite sequencing and chromatin accessibility from DNase I hypersensitivity data. Across multiple human cell lines, EpiBinder consistently outperforms strong sequence-only baselines, improving TF-binding prediction by up to 10% in area under the precision-recall curve. Beyond predictive performance, EpiBinder provides base-level attribution maps that enable systematic interrogation of regulatory context, including candidate methylation-sensitive loci, contextual motif dependencies, and putative TF-TF interactions. These results position EpiBinder as a practical framework for modeling and exploring the local regulatory grammar underlying cell-type-specific TF occupancy.
R. Solozabal, Albert Baichorov, Irina Miodownik et al.· bioRxiv· 0 citations
This review compares convolutional, Transformer-based and graph architectures used to represent local sequence features, chromatin state and three-dimensional genome organisation to their applications to transcription-factor binding, chromatin accessibility, gene expression, non-coding variant prioritisation and regulatory-sequence design.
It is suggested that progress requires reframing seq2func models as continually refined systems, in which targeted perturbation experiments, systematic evaluation and iterative model updates are tightly coupled through artificial intelligence-experiment feedback loops, enabling self-improving models that progressively deepen mechanistic understanding and more reliably support biological discovery.
Masayuki Nagai, A. E. Murphy, Kaeli Rizzo et al.· Nature Genetics· 2 citations
This work developed a scoring-approach for AI-agents to autonomously assess AlphaGenome prediction confidence and accurately differentiate between AlphaGenome’s robust sequence-level recognition across species and its current limitations when interpreting un-fine-mapped regulatory variants.
Priya Ramarao-Milne, Suyu Ma, L. Sng et al.· bioRxiv· 0 citations
Gene expression is governed by regulatory DNA and their associated trans factors acting in specific cell types, yet the sequences underlying this control remain poorly mapped in plants. Genome-pretrained DNA language models provide a route to interrogate regulatory sequence directly, but their attributions have largely been interpreted using bulk or whole-tissue data, and standard attribution pipelines can preferentially highlight sequences downstream of the transcription start (TSS) site rather than promoter-associated signals. Here, we train a celltype-resolved sequence-to-expression model from a single-cell soybean (Glycine max) atlas by coupling a soybean-adapted Genomic Pre-trained Network (GPN) to a shared sequence encoder with 66 cell-type-specific output heads. Across 38,339 protein-coding genes, the model achieves a mean per-cell-type, across-gene Pearson correlation of 0.683 and, recast as a highversus-low expression classification, reaches an area under the ROC curve of 0.92 to 0.97 across tissues, at or above dedicated plant sequence models. We then introduce ContextAware Significance of Cross-gene Attribution for Discovering Elements (CASCADE), a positionspecific statistical framework for identifying model-derived candidate regulatory elements from in silico saturation mutagenesis. Relative to the pooled null used by TF-MoDISco, CASCADE shifts motif recovery from downstream of the transcription start site toward promoter sequence, with 77% of CASCADE-exclusive motifs, compared with 12% of TF-MoDISco-exclusive motifs, falling within the promoter. Applied across the atlas, CASCADE identifies approximately 1.39 million candidate elements spanning broadly active, tissue-restricted and cell-type-restricted classes. Together, these analyses establish a position-aware approach for extracting promoterassociated regulatory hypotheses from sequence models and generate a cell-type-resolved map of candidate cis-regulatory elements.
Ali Farghadan, Robert J. Schmitz, Scott A. Jackson et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.