Jul 2026· Molecular Systems Biology· 0 citations· 20 references
Medicine
Abstract
Large-scale genomic language models (gLMs) hold promise for modeling gene regulation, yet their ability to capture personal gene expression variations remains unresolved. We developed xDecoder, a unified decoding framework that utilizes gLMs and sequence-to-function (S2F) embeddings to learn how personal genetic variation shapes gene expression from paired genome-transcriptome data. Compared to the pretrained genomic models, xDecoder with personalized DNA-RNA training makes cross-individual prediction tractable for seen genes in a few-shot setting. However, zero-shot prediction at unseen loci remains unreliable and gene-dependent, revealing a cross-locus transfer bottleneck of current sequence models. Experiments incorporating individual-level chromatin accessibility suggested that regulatory-state information important for unseen-locus prediction is not fully captured by current DNA-only models. Overall, these results highlight the potential utility of the few-shot setting, the limitations of DNA-only models, and point toward multi-omic, variant-aware frameworks as a promising direction for building personalized regulatory models.
GB.GeneUnet, an 837M-parameter transformer-based U-Net pretrained on 6 trillion tokens from multi-species genomes in OpenGenome2 is introduced, extending genomic context to 1 Mb with up to 100× inference speedup over GeneMoE, a preliminary MoE transformer baseline of similar model size pretrained on the same data.
Ning Sun, William de Vazelhes, Pan Li et al.· bioRxiv· 0 citations
Deep learning models applied to DNA sequences have achieved success in predicting gene expression, chromatin profiles, and variant pathogenicity. However, learning cross-individual differences remains challenging because DNA sequence variation between individuals is small, while gene expression is heavily influenced by non-genetic noise. Prior efforts to predict personalized gene expression from sequence have shown limited generalizability beyond training genes, revealing limitations such as the dilution of variant signals among highly similar input sequences across consecutive convolutional downsampling layers. In this work, we explore architectural modifications addressing these challenges. We propose a contrastive regulatory embedding attention model (CREAM) designed to better capture subtle sequence differences between individuals. To mitigate the impact of non-genetic variability, we decompose gene expression into genetic and non-genetic components and evaluate model performance on the genetic signal. Evaluated on simulated and GTEx transcriptomic datasets, CREAM captures tissue-specific gene expression and significantly outperforms baseline architectures on training genes. CREAM autonomously prioritizes statistically fine-mapped causal expression quantitative trait loci, and introducing an auxiliary L1 inductive bias further sharpens localizing causal variant in unseen genes. However, generalization to predicting expression of unseen genes collapses to near-zero correlation among all tested methods. Single-run model predictive uncertainty capture prediction accuracy in training genes and mirrors cross-run consistency in test genes. While accurate inference on unseen genes remains an open problem, our results highlight key obstacles and suggest directions for modeling personalized gene regulation from DNA sequence.
Zhirui Hu, Jason Ku, Katherine S. Pollard· bioRxiv· 0 citations
Genomic Foundation Models (GFMs) are increasingly used for large-scale sequence analysis and generation. Compared with frontier language models, GFMs are typically smaller and frequently operate on long genomic sequences, with evaluation often requiring preservation of biologically meaningful structure and sequence-level relationships. Although low-precision post-training quantization (PTQ) has shown substantial memory and throughput benefits for general-purpose language models, it remains unclear whether these benefits transfer to GFMs given their distinct model scales, sequence characteristics, and evaluation requirements. We present an empirical case study of FP8 post-training quantization applied to GenomeOcean, a computationally efficient genomic foundation model with strong reported performance across diverse genomics tasks [Zhou et al., 2025]. Its range of model scales, from 100M to 4B parameters, provides a useful setting for examining how quantization effects vary with model size. We evaluate FP8 across two primary GFM inference regimes—embedding extraction and autoregressive generation—and assess its impact along two dimensions: biological fidelity relative to BF16 baselines and system-level efficiency in terms of throughput, memory usage, and energy efficiency. We find that FP8 largely preserves biological fidelity across the evaluated scales and inference regimes, while reducing GPU memory footprint at 4B scale and improving energy efficiency during autoregressive generation. However, realized throughput gains remain substantially below FP8’s theoretical 2× hardware ceiling, with a best-case improvement of 19.3% in autoregressive generation and benefits varying strongly by model scale and workload. Autoregressive generation shows the clearest gains, driven largely by KV-cache compression, whereas embedding extraction provides limited or negative throughput benefits at smaller model scales. We attribute this theory–practice gap to the interaction of model-scale effects, memory-system bottlenecks, and software-stack limitations. These findings highlight the need for workload-specific empirical evaluation before adopting low-precision inference in scientific foundation models. Code availability https://github.com/jgi-genomeocean/genomeocean_efficiency
Mutian Yu, Robert Egan, Fengchen Liu et al.· bioRxiv· 0 citations
It is suggested that progress requires reframing seq2func models as continually refined systems, in which targeted perturbation experiments, systematic evaluation and iterative model updates are tightly coupled through artificial intelligence-experiment feedback loops, enabling self-improving models that progressively deepen mechanistic understanding and more reliably support biological discovery.
Masayuki Nagai, A. E. Murphy, Kaeli Rizzo et al.· Nature Genetics· 2 citations
A representation-accessibility analysis of frozen genomic language models across regulatory, epigenetic, promoter, splice-site, and variant-effect prediction tasks shows that local biological signal is partially present in frozen representations, but is not always accessible through final pooled embeddings.
Nirjhor Datta, Swakkhar Shatabda, M. S. Rahman· 0 citations
Sequence-to-function models learn regulatory features from genomic sequence, but they remain limited in their ability to predict gene-expression differences among individuals. Cell-type-specific regulatory effects may be obscured in bulk RNA sequencing, whereas paired genotype and single-cell expression cohorts remain small. We evaluated whether deconvolution of bulk RNA-seq could provide scalable cell-type-specific targets for personal-genome expression prediction. GTEx v8 bulk RNA-seq from six tissues was deconvolved with BayesPrism using single-nucleus reference profiles, producing targets across 83 tissue–cell-type contexts. Deconvolved expression agreed with matched pseudobulked GTEx single-nucleus RNA-seq, with median donor-level Pearson correlations across genes ranging from 0.53 to 0.73 by tissue. We compared genotype-feature models, regressors trained on frozen Enformer representations, and fine-tuned Enformer and Borzoi models. Across random and nonlinear-enriched gene sets, sequence-derived approaches generally outperformed genotype-feature baselines, while frozen Enformer features were competitive with end-to-end fine-tuning. For the random gene set, Fisher-averaged Pearson correlations were 0.122–0.142 for sequence-derived approaches and 0.081–0.086 for genotype-feature baselines in a coverage-aware sensitivity analysis. Model performance was positively associated with deconvolution–pseudobulk agreement for sequence-derived models (r = 0.35–0.43 across tissue–cell-type contexts), suggesting that target reliability may constrain downstream prediction. Context-specific Enformer fine-tuning did not materially out-perform a shared, combined-context strategy. These results support deconvolution as a feasible approach for generating cell-type-resolved training targets, while showing that target quality and limited cohort size remain important constraints. Frozen pretrained representations provide a computationally efficient and competitive baseline for personal sequence-to-expression modeling.
Stephen Sim, Li Shen· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.