Skip to content
Open access

A self-supervised DNA foundation model with collapse-resistant multimodal fusion

Aug 2026 · bioRxiv · 0 citations · 10 references
Biology

TL;DR

This work presents a self-supervised DNA-centric multimodal foundation model integrating DNA sequence embeddings with local and global chromatin accessibility in a shared encoder to produce reusable window-level embeddings that improve regulatory activity prediction, regulatory signal ranking and chromatin accessibility peak detection.

Abstract

Genomic foundation models pretrained on DNA sequence have achieved strong performance across many tasks, but sequence-only representations cannot fully capture regulatory information from additional DNA-centric modalities. Existing multimodal genomic models are optimized for specific prediction tasks rather than reusable embeddings. Directly fusing heterogeneous modalities is challenging because sparse, peak-shaped regulatory signals and dense sequence embeddings have markedly different statistical structures, making naive alignment prone to near-zero solutions. We present a self-supervised DNA-centric multimodal foundation model integrating DNA sequence embeddings with local and global chromatin accessibility in a shared encoder to produce reusable window-level embeddings. We show that global normalization alleviates this collapse, enabling effective joint learning. The resulting embeddings improve regulatory activity prediction, regulatory signal ranking and chromatin accessibility peak detection, achieving a 4.6-fold AUPRC improvement over the DNA-only baseline, with further gains on external ClinVar, GTEx eQTL and PBMC caQTL datasets.

Read PDF

Similar papers

Open access Jul 2026

Deep Learning Predicts Dissimilar DNA-DNA Binding and Engineers Hyperconnected Networks

BINND is developed, a binding and interaction neural network to predict non-orthogonal DNA interactions, which could aid diagnostic, bioengineering, and DNA origami design, and supporting a shift toward exploiting the full sequence space.

Karishma Matange, Gunavaran Brihadiswaran, Kyle J. Tomek et al. · 0 citations
Open access Aug 2026

Multi-task adversarial autoencoder for functional genomic element generation with preserved biophysical properties

Generative modeling of genomic sequences presents a stringent test for deep learning, requiring the capture of long-range dependencies and functional constraints beyond local nucleotide statistics. Existing architectures frequently collapse to limited modes or reproduce shallow nucleotide distributions without encoding functional semantics. We introduce the Multi-Task Adversarial Autoencoder (MT-AAE), a hybrid generative framework that integrates adversarial regularization with auxiliary functional and biophysical objectives to enforce structured latent representations. Evaluated on an empirical human gene corpus, MT-AAE achieved a Train-on-Synthetic-Test-on-Real (TRTS) accuracy of 74.7%, compared with 41.0% for a standard GAN baseline. Stratified analysis further showed that functional discriminability increased to 89.3% when sequence lengths aligned with the model’s architectural window. Importantly, the learned representations exhibited emergent biological structure: synthetic sequences spontaneously preserved cis -regulatory syntax, including canonical TATA-box motifs recovered across 100% of generated promoter sequences without explicit rule encoding, though positional placement relative to the TSS was not statistically significant (KS $$p=0.90$$ ), and the high occurrence rate is partly attributable to the AT-rich composition of the generated sequences. Representation-level validation using frozen DNABERT-2 and DNABERT-S embeddings confirmed that the generated sequences retained functional information beyond shallow k-mer statistics. Cross-species evaluation on Mus musculus sequences further demonstrated species-specific learning consistent with known human–mouse regulatory divergence. The framework also mitigated mode collapse, maintaining near-uniform generation across functional classes ( $$R_g \approx 1.0$$ ), including rare categories such as tRNAs ( $$ < 2\%$$ of the dataset). These findings position MT-AAE as an effective framework for biologically constrained genomic sequence generation.

Shamsuddeen Adamu, H. Alhussian, S. Abdulkadir et al. · 0 citations
Open access Aug 2026

Structure-agnostic protein–ligand binding affinity prediction via hierarchical representation alignment

Abstract Motivation To enable real-world protein-ligand affinity prediction, not only out-of-distribution generalization but also robustness to variable structural availability and quality should be considered in model design. Results We present AlignNet, a hierarchical representation alignment framework that mitigates intra- and inter-molecular heterogeneity to learn robust protein-ligand embeddings for generalizable affinity prediction, even from sequence-level inputs. Its intra-molecular module projects unimodal and multimodal features into a unified space, aligning augmented multimodal views for feature fusion and unimodal with multimodal embeddings to distill multimodal priors for structure-agnostic inference. Its inter-molecular module aligns protein and ligand embeddings for cross-molecular integration. Extensive experiments show that AlignNet (i) achieves highly competitive performance, with up to a 20.4% gain in SCC on the challenging LBA 30% split under sequence-only settings, suggesting improved out-of-distribution generalization; and (ii) learns well-separated affinity-related clusters, supporting reliable structure-independent prediction. Availability and implementation AlignNet is available at https://github.com/altriavin/AlignNet.

Xiaowen Hu, Hongyi Huang, Hao Sun et al. · 0 citations
Open access Jul 2026

xDecoder unlocks the potential of genomic foundation models for few-shot personal gene expression prediction.

Large-scale genomic language models (gLMs) hold promise for modeling gene regulation, yet their ability to capture personal gene expression variations remains unresolved. We developed xDecoder, a unified decoding framework that utilizes gLMs and sequence-to-function (S2F) embeddings to learn how personal genetic variation shapes gene expression from paired genome-transcriptome data. Compared to the pretrained genomic models, xDecoder with personalized DNA-RNA training makes cross-individual prediction tractable for seen genes in a few-shot setting. However, zero-shot prediction at unseen loci remains unreliable and gene-dependent, revealing a cross-locus transfer bottleneck of current sequence models. Experiments incorporating individual-level chromatin accessibility suggested that regulatory-state information important for unseen-locus prediction is not fully captured by current DNA-only models. Overall, these results highlight the potential utility of the few-shot setting, the limitations of DNA-only models, and point toward multi-omic, variant-aware frameworks as a promising direction for building personalized regulatory models.

Shumin Li, Ruibang Luo, Yuanhua Huang · 0 citations
Open access Aug 2026

ProtEnrich: Residual Multimodal Enrichment of Protein Sequence Embeddings

Over eight diverse protein foundational models trained on 550,120 SwissProt proteins with AlphaFold structures, enriched embeddings improved zero-shot remote homology retrieval, increasing Precision@10 and MRR by up to 0.13 and 0.11, respectively.

Gabriel Bianchin de Oliveira, Fahad Saeed · 0 citations
#protein folding Preprint Aug 2026

Off-Manifold Collapse in Guided Protein Language Models

A cheap density prior is introduced over natural protein activations and keeps only the candidates that remain typical under it, a training-free post-hoc step the authors call Mahalanobis filtering that improves both the property score and the structural plausibility of the sequences it keeps at negligible cost, without touching the generator, and transfers across different guidance methods.

Shuibai Zhang, Xin-Chi Liu, Fred Zhangzhi Peng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.