Skip to content
Open access

SAR11 Genome Atlas: a genome and gene catalog for functional profiling of the most abundant bacterial clade in the ocean

Aug 2026 · bioRxiv · 0 citations
Biology

TL;DR

The SAR11 Genome Atlas is presented, an interactive ortholog group (OG)-centered web resource that integrates 542 SAR11 genomes, including all 132 cultured strain genomes, with functional annotations, synteny, phylogenetic distribution, metatranscriptomic expression, and predicted protein structure information.

Abstract

The SAR11 clade, also known as the order Candidatus Pelagibacterales, is among the most abundant bacterial lineages in the ocean and plays central roles in marine biogeochemical cycles. However, many SAR11 genes remain functionally uncharacterized, highlighting the need for a comprehensive, integrated catalog that supports genomic, functional, and ecological analyses across the clade. Here, we present the SAR11 Genome Atlas, an interactive ortholog group (OG)-centered web resource that integrates 542 SAR11 genomes, including all 132 cultured strain genomes, with functional annotations, synteny, phylogenetic distribution, metatranscriptomic expression, and predicted protein structure information. To demonstrate its utility, we used environmental expression profiles to identify OGs associated with high-latitude environments, recovering OGs known to be involved in cold adaptation and proposing a hypothesis for the function of uncharacterized protein. We further analyzed phylogenetic distribution patterns to identify mutually exclusive functional modules, including candidate alternative systems for Mn/Zn homeostasis and phosphate acquisition, and to associate these modules with distinct oceanographic environments. Together, these case studies demonstrate that the SAR11 Genome Atlas supports complementary analyses that connect environmental signals to genes of interest and use phylogenetic or functional distributions to generate hypotheses about ecological specialization. Through a user-friendly web interface, the SAR11 Genome Atlas enables researchers to explore genomic, environmental, and structural information without specialized computational expertise. All data and analysis outputs are freely accessible online at [https://stsnsn.github.io/SAR11_Atlas/]. The SAR11 Genome Atlas thus provides a scalable framework for generating and testing hypotheses that connect SAR11 genomic variation to protein function and oceanographic context, supporting advances in marine microbial ecology and biogeochemistry.

Read PDF

Similar papers

Open access Aug 2026

ChlORIS: Chloroplast Orthologs Resource & Identification Suite

Chloroplast or plastid genomes are essential resources for studying the evolution and diversity of algae and land plants. Although thousands of plastid genomes have been sequenced, their full potential has not been realised; derived resources such as orthogroup databases and reference datasets for metagenomic profiling remain underdeveloped. We present the ChlORIS database to address these problems across all algal phyla. From 2,254 publicly available algal plastid genomes, after dereplication we clustered 2,531 orthogroups from the annotated proteins and selected 496 orthogroups with consistent gene naming, enabling cross-genome comparisons of homologous plastid proteins. We further selected 224 core orthogroups, each containing more than 10 protein sequences, for which we produced score-calibrated hidden Markov models (HMMs), multiple sequence alignments and predicted protein structures. The value of these resources for phylogenomics is demonstrated through a large-scale plastid phylogeny of 859 taxa spanning all major algal lineages. We characterised the protein HMMs by cross-referencing them to Pfam domains and calibrated score cutoffs for reliable detection. The metagenomic database, HMM library, nucleotide and amino acid alignments, predicted structures and protein metadata, cross-linked to UniProt and InterPro (Pfam), are openly available on the ChlORIS website at https://chloris.codeberg.page/.

Yuhao Tong, Vanessa Rossetto Marcelino, Robert Turnbull et al. · 0 citations
Open access Aug 2026

GenomeCompendium: A database for the integrated analysis of repeats, assembly quality and functional content of complete prokaryotic genomes

The GenomeCompendium is released, a public database and interactive analysis tool for complete prokaryotic genomes and it is shown that complex, repeat-rich genomes are more common than previously estimated.

Tiberiu Totu, Garance Jaques, B. Heiniger et al. · 0 citations
Review Open access Aug 2026

Genome-scale insights into the metabolic landscape and evolutionary development of Bifidobacterium bifidum

This study reconstructed the first comprehensive pangenome of B. bifidum using 1,351 high-quality genomes, including metagenome-assembled genomes to identify species-specific genetic and functional features and identified significant gain-of-function events.

Emanuele Selleri, G. Longhi, C. Tarracchini et al. · 0 citations
Open access Jul 2026

Phylogenize2: robust phylogenetic methods link genes to phenotypes across host-associated and environmental microbiomes

In microbiome studies, associations between microbial functions and the environment are often confounded by phylogeny. While some methods explicitly account for this confounder, they require information about genome content, limiting their use in biomes where few genomes have been available. To make these methods more universally accessible, we have developed Phylogenize2, a redesigned phylogeny-aware tool for linking microbial gene families to abundance phenotypes. Phylogenize2 integrates large metagenome-assembled genome collections, including both biome-specific collections from MGnify and a broadly sampled general purpose database, GlobDB, to substantially expand species coverage, allowing its application in environments like the mouse gut and ocean. In addition, by default, Phylogenize2 uses a new robust phylogenetic testing framework that has been optimized for microbial abundance data, while also allowing the use of other comparative methods such as POMS. In an experimental mouse study, Phylogenize2 identifies that Muribaculaceae with higher abundance on a high-fat diet are enriched for proteins in the thioredoxin family, with likely roles in oxidative stress. When we apply Phylogenize2 to a polar ocean study, we find that a molybdenum-dependent PaoABC/YagTSR-like aldehyde oxidoreductase system differentiates mesopelagic from surface-dwelling Flavobacteriaceae, suggesting that aldehyde detoxification may be important for organisms that degrade marine snow. Together, these results show that Phylogenize2 expands phylogeny-aware microbiome analysis beyond the human gut and can provide insight into the genetic basis of microbiome-encoded traits in diverse environments. Importance Microbiome studies often set out to identify which microbes are more or less abundant across environments, but these patterns can be difficult to interpret. Phylogenize2 is an open-source software package that allows researchers to ask whether individual microbial gene families are associated with the environment across independent branches of the microbial tree of life. By incorporating large collections of genomes from uncultivated microbes, as well as modern statistical methods designed for microbial abundance data, Phylogenize2 makes this approach practical for microbiomes beyond the human gut, including in model organisms like lab mice and free-living environments like the ocean. We also provide a pipeline that allows the use of new genome collections. In two case studies, we demonstrate that Phylogenize2 effectively prioritizes specific genes and pathways from metagenomic data, thereby leading researchers from changes in microbial abundance to more biologically interpretable explanations.

Kathryn Kananen, Nghi Tran, Patrick H. Bradley · 0 citations
Aug 2026

A pan-genome perspective uncovers the core genetic basis and evolutionary adaptation of lipid synthesis and vesicular transport in Nannochloropsis.

A pan-genome of 17 Nannochloropsis species comprising 14,851 gene families is constructed and a distinct genetic architecture for lipid metabolism is defined: Gene families associated with vesicular transport formed a conserved core functional module, whereas the genetic collection for lipid metabolism showed greater plasticity and was primarily classified as part of the soft-core genome.

Pengjuan Zhang, Lijun Miao, Hua Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.