Skip to content
Preprint

BioM-JEPA: joint-embedding prediction of graph-connected gene blocks in single cells

Aug 2026 · 0 citations · 23 references
Computer Science

TL;DR

Results support graph-connected gene blocks as useful prediction units for JEPA-style representation learning in single-cell biology by supporting block-level prediction of graph-connected gene blocks defined by protein-association and corpus-derived coexpression evidence.

Abstract

Single-cell transcriptomes are sparse observations of coordinated biological programmes, yet most self-supervised models learn by reconstructing individual genes. Here we present BioM-JEPA, a joint-embedding predictive architecture that instead predicts aggregate representations of graph-connected gene blocks defined by protein-association and corpus-derived coexpression evidence. A student network infers each target-block representation from the remaining genes in a cell, while a slowly updated teacher supplies the corresponding target from the full observed gene set. Under the reported extraction procedure, block-level prediction produced embeddings with higher effective rank and weaker association with detected-gene depth in the tested diagnostics than token-prediction, random-block and reconstruction controls. Across CellBench tasks, frozen BioM-JEPA embeddings retained expression, pathway and neighbourhood information and achieved the lowest aggregate perturbation-response error among the evaluated models. Representation diagnostics were also consistent with canonical pancreatic programmes and compositional relationships between genetic perturbations. Linear attention avoids constructing a quadratic gene-by-gene attention matrix; in a matched one-epoch hPancreas experiment at batch size 8, BioM-JEPA provided 5.75-fold higher fine-tuning throughput and 3.76-fold higher held-out embedding throughput than scFoundation. Together, these results support graph-connected gene blocks as useful prediction units for JEPA-style representation learning in single-cell biology.

View source

Similar papers

Open access Jul 2026

FloREN: Decoding Immune Regulatory Networks through Interpretable Graph Transformer Patient Representations

A Framework for Learning Over REgulatory-Embedding Networks (FloREN), a supervised and interpretable sample representation method that enables improved sample stratification and biomarker discovery and supports downstream analyses that found specific immune network mechanisms in immune-mediated inflammatory diseases (IMIDs).

Iñigo Clemente‐Larramendi, S. Hillion, D. Cornec et al. · 0 citations
Preprint Aug 2026

bioMoR: Biology-Guided Mixture-of-Recursions for Effective Genomic Learning

This work proposes bioMoR, which is the first framework to apply MoR to gene-level and pathway-level learning, and identifies three locations for integrating structured biological knowledge within an MoR backbone: graph-based information sharing refines token embeddings, a structural bias guides self-attention toward biologically related tokens, and a graph-aware router uses neighborhood information to determine each token's recursion depth.

Koushik Howlader, Tirtho Roy, Md Tauhidul Islam et al. · 0 citations
Open access Jul 2026

Coding agents author interpretable single-cell embedding models from the literature

The single-cell literature catalogs cell states as validated marker-gene programs — a sparse, compositional prior. Conventional embedding methods do not leverage this prior and learn cell-state structure de novo from the expression matrix, producing dense dimensions needing post-hoc interpretation and batch correction. Here we show coding agents can author single-cell embedding models directly from the literature. Given a scenario that focuses this literature lens on a chosen biological subdomain, the agent edits a structured Python template, curating named, literature-cited gene programs and composing them into axes, without a gene-set database, training, or sight of the data. Across mouse and human tissues these zero-shot embeddings are competitive in biological quality with conventional, foundation-model, and program-informed baselines, batch-robust by construction and reproducible across runs, complementing data-driven embeddings. Because each dimension is a named, cited gene program, the embedding is interpretable and auditable, and its composable axes can be steered into a developmental tree.

Niklas Brunn, S. M. Krißmer, Maximilian Frosch et al. · 0 citations
Open access Jul 2026

scGraphVerse: a modular workflow for single-cell gene network inference

Abstract Motivation Inferring gene networks from single-cell RNA sequencing data is challenging due to high sparsity, dimensionality, and technical noise. Current pipelines lack the multi-dataset integration and comprehensive post-processing analysis. Results scGraphVerse is an R package that integrates multiple algorithms (GENIE3, GRNBoost2, ZILGM, PCzinb, and JRF) with extensive evaluation and visualization tools. Its modular workflow supports early, late, and joint integration strategies for multi-dataset analysis, providing standardized input/output interfaces and biological interpretation tools, including community detection, pathway enrichment, and literature mining. Benchmarking on simulated data showed model-based methods (PCzinb and ZILGM) perform well with limited sample sizes, while JRF performs best as the network size and dataset numbers increase. A PBMC case study demonstrates JRF’s ability to identify literature-supported regulatory communities across donors. Availability and implementation The package is available in Bioconductor 3.22 at https://bioconductor.org/packages/release/bioc/html/scGraphVerse.html. Code and examples: https://github.com/ngsFC/scGV_analysis.

Francesco Cecere, D. De Canditiis, Annamaria Carissimo et al. · 0 citations
Preprint Aug 2026

CellWorld: From Gene-Level Reconstruction to Latent Cell Prediction in Spatial Transcriptomics Foundation Models

This paper shows that latent-space predictive pretraining can provide a scalable route to foundation models for spatial transcriptomics. Existing spatial transcriptomics foundation models primarily reconstruct masked gene identities or expression values, potentially encouraging the reproduction of assay-specific technical variation and limiting representation transferability. To avoid directly reconstructing such variation, we shift the prediction target from observed gene measurements to latent cell representations and introduce CellWorld, which predicts the latent representations of masked cells from visible spatial context and a limited partial-expression hint. We pretrain four CellWorld variants, spanning 5.74M to 94.56M trainable parameters, on a corpus of 46 million human cells. Our controlled scaling experiments show that performance improves with model capacity, particularly on spatial tasks, while spatial transfer depends more on sufficient optimization and broad biological source diversity than on cell count alone. Across four held-out datasets, even CellWorld-Small, with 5.74M trainable parameters, outperforms every baseline on all 11 linear-probe benchmarks and all seven fine-tuned spatial benchmarks. Most notably, a frozen CellWorld-Large pretrained on only 5\% of the corpus with broad biological source coverage outperforms every fully fine-tuned baseline across all seven spatial benchmarks. Code is available at https://github.com/UoM-HealthAI/CellWorld.

Haiping Liu, Qian Zhao, Lijing Lin et al. · 0 citations
Open access Aug 2026

SLIM: A small linear model with STRING embeddings for single-cell genetic perturbation prediction

Predicting cellular responses to genetic perturbations is central to understanding gene function and prioritizing therapeutic targets, but experimental screens cannot exhaustively cover genes, cell types, and perturbation combinations. Recent benchmarks have shown that simple baselines can match or outperform substantially more complex models, suggesting that informative biological priors may be as important as model capacity. Here we present SLIM, a lightweight extension of the bilinear model of Ahlmann-Eltze et al. SLIM represents perturbations with 64-dimensional embeddings derived from the STRING protein network and predicts mean transcriptional responses through a closed-form ridge-regression estimator. It then constructs single-cell populations by retrieving training cells and rescaling each gene to match the predicted mean. We evaluated SLIM against four deep learning models and two simple baselines on four single-gene perturbation datasets and one combinatorial perturbation dataset. Across these within-dataset benchmarks, SLIM achieved competitive mean-response accuracy, ranked first in eight of twelve single-gene dataset–metric comparisons, and produced substantially lower maximum mean discrepancy values than the evaluated alternatives. The model has 640 trainable parameters and fitted each benchmark dataset in under 10 seconds on a CPU. These results show that compact biological representations can support accurate and computationally efficient perturbation prediction. Code is available at https://github.com/RasmussenLab/SLIM. Key Points SLIM combines a closed-form bilinear predictor with STRING-derived perturbation embeddings. Across five within-dataset benchmarks, SLIM achieved competitive mean-response prediction with only 640 trainable parameters. SLIM builds cell populations by rescaling retrieved training cells to the predicted mean, so they inherit realistic cell-to-cell variation and gene–gene covariation. The results highlight the importance of perturbation representations and population-construction procedures in low-data benchmarks. SLIM fits each benchmark dataset in under 10 seconds on a standard CPU.

Dewei Hu, M. Pielies Avellí, L. J. Jensen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.