Integrating single-cell RNA-sequencing (scRNA-seq) data across species is hindered by evolutionary divergence, technical batch effects, and the reliance on one-to-one orthologs. Here, we present Unify, a transfer learning methodology that learns universal cell embeddings by defining functionally coherent, multi-modal macrogenes. This is achieved by combining RNA expression with embeddings from protein language models and general-purpose language models. Unify transcends species boundaries, enabling cross-species comparisons beyond strict gene-level homology. Unify corrects batch effects while preserving conserved biological signals across vast evolutionary distances and enables more accurate prediction of perturbation responses across species, such as from mouse to human. Applied to species separated by over 700 million years, Unify reconstructs more accurate multi-species cell-type evolutionary trees and uncovers convergent gene programs. Together, these results establish Unify as a powerful method for comparative single-cell genomics and evolutionary biology. Integrating single-cell RNA-sequencing (scRNA-seq) data across species is still technically challenging. Here, the authors report a transfer learning framework designed to integrate scRNA-seq data across species by combining RNA expression with embeddings from protein language models and general-purpose language models.
Hua-Wen Zhong, Wenkai Han, Guoxin Cui et al.· Nature Communications· 0 citations
ABSTRACT Earth's biodiversity is central to ecosystem health and resilience, providing essential functions and services. The Red Sea is a recognised marine biodiversity hotspot with high endemism and unique environmental conditions that support extensive but poorly resolved biodiversity. Here, we applied metagenomic analyses to sediment samples collected from coastal to deep‐sea environments during the Red Sea Decade Expedition 2022 to characterise biodiversity across the web of life. From a single shotgun assay per sample, this approach simultaneously characterised the sediment microbiome, which amplicon‐based surveys recover only through parallel, targeted assays, and extended detection to higher eukaryotes. Using high‐throughput sequencing, we generated 12.8 billion sequences, revealing taxa covering all domains of life. Although eukaryotic sequences represented only 0.7% of the taxonomically annotated dataset, we managed to identify 679 eukaryotic families. Prokaryotic diversity was high, as expected in a basin‐scale sampling coupled with high sequencing depth, with groups covering a wide functional array. Community structure analyses revealed depth‐driven stratification of open‐ocean benthic microbial communities and latitudinal structuring of coastal benthic eukaryotes. Overall, this dataset provides an empirical reliability–coverage trade‐off with direct consequences for the design of eDNA monitoring programmes targeting conservation‐priority taxa, and clear priorities for taxa specific reference‐database expansion.
Elisa Laiolo, Christopher A. Hempel, Balegh A. Abukabbos et al.· Environmental Microbiology· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.