Bacterial small RNAs (sRNAs) regulate gene expression by base pairing with target mRNAs, yet transcriptome-wide interactome mapping has shown that many sRNA–mRNA interactions detected in vivo have modest or no regulatory effect using orthogonal reporter assays. The features that determine functional outcome remain poorly defined. Here, we integrated Hfq-CLASH interactome mapping with matched transcriptomic and proteomic profiling in Escherichia coli and developed an interpretable machine-learning framework to identify the determinants that distinguish functional from non-functional interactions. Using sequence, structural, thermodynamic, duplex and protein-occupancy features, transcriptomic and proteomic responses were predicted with above-chance performance, achieving AUCs of 0.78 and 0.74, respectively. Feature attribution revealed that physical pairing alone is insufficient for regulation; instead, regulatory outcome is shaped by a coordinated interplay between RNA secondary structure, thermodynamic accessibility and local protein-binding context. Target-side Hfq occupancy emerged as a positive predictor of functional regulation, whereas AR2-domain occupancy on the sRNA was associated with non-responsive interactions, suggesting that distinct ribonucleoprotein states may separate productive regulation from non-productive binding. These findings indicate that the regulatory fate of an sRNA–mRNA interaction is an emergent property of its biophysical context and protein-binding environment, rather than a direct consequence of physical pairing alone. GRAPHICAL ABSTRACT
F. Safari, Daniel G. Mediati, Saleh Alquethamy et al.· bioRxiv· 0 citations
Single-cell and spatial foundation models promise transferable biological representations, yet their generality remains largely untested across modalities, biological domains and analytical tasks. We benchmarked six representative models, Nicheformer, CellPLM, scGPT-spatial, GenePT, scELMo and Novae, using a harmonised framework spanning scRNA-seq, spatial transcriptomics and Perturb-seq. We evaluated zero-shot and continually pretrained clustering, supervised annotation, marker-gene concordance and perturbation prediction. Model performance was strongly conditional: expression-trained cell-level transformers best resolved many cell-identity tasks, spatial and graph-aware models better preserved tissue architecture, and language-derived gene embeddings were competitive for selected perturbation-response metrics. No model dominated across tasks, and rankings shifted with modality, preprocessing, tokenisation, biological prior, domain shift and metric choice. This benchmark provides practical guidance for model selection and argues that future models should be judged by biological generalisation, interpretability and perturbation-grounded validity, not by scale or leaderboard performance alone.
Sally Chen, Roxana Zahedi, Lucy Chhuo et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.