Back to feed
Open access

Annotation of glycoside hydrolases in unassembled metagenomes using CAZyOGH

Jul 2026 · Bioinformatics Advances · Vol 6 · 0 citations · 44 references
Medicine

Abstract

Abstract Motivation Functional characterization of microbiomes often relies on the sequencing of metagenomic DNA extracted from environmental samples, with current approaches using metagenome-assembled genomes (MAGs). Although glycoside hydrolases (GHs) are central to carbon cycling, accurate annotation of GHs in metagenomic datasets remains challenging due to the multidomain architecture of carbohydrate-active enzymes and the prevalence of unassembled short reads due to limitations in the MAG-generation process. Results Here, we present CAZyOGH (CAZymes Open-source GH annotation), a curated reference database for the domain-specific identification of 135 protein domains spanning 99 GH families with well-defined catalytic domain signatures. CAZyOGH focuses on individual GH domains, enabling robust annotation of both assembled and unassembled metagenomic data. We validated CAZyOGH by reanalyzing genomes listed in CAZy db, where predicted GH profiles closely matched reported values. Next, we used CAZyOGH to analyze 12 human gut metagenomes and 12 newly sequenced soil microbiomes to reveal environment-specific GH repertoires. By accurately detecting catalytic domains independent of the genomic context, CAZyOGH improves sensitivity and specificity in short-read metagenomic annotation. This framework provides a scalable and reproducible approach to investigate carbohydrate-active enzymes across ecosystems, advancing our capacity to characterize microbial functional potential in global carbon cycling. Availability and implementation CAZyOGH data is available on figshare (https://figshare.com/projects/CAZyO_GH/267770).

Read PDF

Similar papers

Open access 2025

Metagenomic and Bioinformatics Analysis of Sebacina vermifera for Next-Generation Biofertilizer Development

Background: Sebacina vermifera is a fungus that belongs to the Basidiomycota phylum (order Sebacinales). It has potential as a biofertilizer because it forms mutualistic relationships with many plant species, including orchids and other flowering plants. However, it is still not fully understood how this fungus plays a role in the soil/rhizosphere ecosystems where it occurs. Methods: To profile S. vermifera-associated microbial community structures in relation to the different agroecological zone types, a multi-platform metagenomic sequencing approach was employed using sequencing technology platforms such as PacBio long read, Illumina short Read and Oxford Nanopore. Data processing for these metagenomic sequence assemblies included multiple steps including quality control using Trimmomatic and fastp, metagenome assembly with MEGAHIT and SPAdes, taxonomic profiling with Kraken2 and MetaPhlAn4, and functional annotation through EggNOG-mapper, KEGG Orthology, and CAZy databases. Network and comparative genomic analyses were also performed to characterise potential microbial interactions, as well as unique gene content. Results: The results of metagenomic analyses showed that genes associated with phosphate solubilization (e.g., phytases, acid phosphatases), nitrogen fixation (e.g., nifH, nifD), production of siderophores, and the biosynthesis of indole-3-acetic acid were present. The association of S. vermifera with rhizosphere microbial networks increased the occurrence of interactions between nitrogen-fixing bacteria, arbuscular mycorrhizal fungi, and plant growth-promoting rhizobacteria. Unique effector proteins and secreted hydrolases were identified that were distinct from those of related fungal species. The field trials demonstrated a 34-42% increase in plant biomass, a 28% increase in phosphorus uptake, and a 19% decrease in applied chemical fertilizer. Conclusion: With its rich repertoire of functional genes and beneficial interactions with other microorganisms, Sebacina vermifera represents a potential new source of biofertilizers for use in agriculture. The use of this fungus will result in greater crop yields, less dependence on chemical fertilizers, and healthier soils, thereby supporting the development of sustainable and climate-resilient agricultural systems.

Prasun Craven · 0 citations
Jul 2026

Multidomain annotation of carbohydrate-active enzymes beyond CAZy domains with GeneHunt2.

Carbohydrate-active enzymes (CAZymes) are central to carbohydrate metabolism, yet their functional annotation is typically restricted to catalytic CAZy domains, overlooking the broader multidomain architectures in which these domains operate. Here, I present GeneHunt2, a scalable framework for multidomain annotation of CAZymes that integrates curated HMM profiles from dbCAN and Pfam into a unified, deduplicated database, enabling systematic identification of both CAZy and non-CAZy domains. After confirming the robust recovery of CAZy domain assignments using GeneHunt2, I investigated the detailed multidomain architecture of over 3.75 million CAZyme sequences: more than 40% were multidomain, and non-CAZy partner domains constituted a substantial fraction of detected partner domains. Using a quantitative framework that combines co-occurrence enrichment, domain adjacency, positional bias, and partner-specificity scoring, I next distinguished family-specific modules, auxiliary domains, and promiscuous partners. This approach recapitulates known CAZyme-CBM relationships and extends beyond CAZy definitions by identifying numerous Pfam domains including many domains of unknown function (DUFs) that are specifically and non-randomly associated with particular CAZy families. By enabling reproducible, multidomain-aware annotation, GeneHunt2 facilitates data-driven hypotheses about poorly characterized domains and widens the functional interpretation of carbohydrate-active proteins beyond their catalytic cores.

R. Berlemont · 0 citations
Review Open access Jul 2026

nf-core/magmap: Map metatranscriptomes to large collections of genomes

Abstract Summary The lack of publicly available reference genomes has forced annotation of metatranscriptomes to either use direct alignment of sequence reads to reference databases or de novo assembly. As more and more natural environments are covered by metagenomic surveys, this is rapidly changing. This opens up the possibility of genome-resolved studies of prokaryotic metatranscriptomes by mapping to genomes from public repositories or metagenome-assembled genomes derived from the same environment. Here, we present the nf-core/magmap pipeline that provides a reproducible, easy-to-access, and well-documented workflow for selecting reference genomes, mapping to them, and quantifying features. Genomes can be drawn from public sources or originate from private collections. The pipeline is primarily aimed at prokaryotic communities but can, together with collections of reference mature gene sequences, also be applied to eukaryotes. Availability and implementation The nf-core/magmap pipeline is implemented in Nextflow and part of the nf-core collaboration. The pipeline is available at the nf-core website (https://nf-co.re/magmap) and GitHub (https://github.com/nf-core/magmap).

Danilo Di Leo, E. Nilsson, George Westmeijer et al. · 0 citations
Aug 2026

Large language models enhance annotation of enzymes in metagenomes

Metagenomic data have notable biological potential, but their functional interpretation is frequently impeded by incomplete protein function annotations. Accurate enzyme annotation is essential for elucidating the metabolic capabilities of microbial communities within metagenomic datasets. To address this challenge, we developed FEDKEA, an enzyme annotation tool leveraging protein language models, and provided a web platform for its use. In addition, we designed a user-friendly, FEDKEA-based metagenomic pipeline, MEnzMap, which encompasses the entire analysis workflow—from raw data quality control to function prediction and downstream analyses. Applying MEnzMap to human gut metagenomic data from the iHMP2 project, we generated a comprehensive enzyme profile landscape for both healthy individuals and patients with inflammatory bowel diseases. These tools provide an efficient method for the functional annotation of microbial dark matter and facilitate the identification of disease-associated enzymes.

Lei Zheng, Bowen Li, Siqi Xu et al. · 0 citations
Review Open access Aug 2026

Genome-scale insights into the metabolic landscape and evolutionary development of Bifidobacterium bifidum

Background: Bifidobacterium bifidum (B. bifidum) is an infant gut symbiont specialized in degrading host-derived glycans. Despite its relevance in early life, the species’ genomic diversity has not yet been comprehensively surveyed, and current reference collections capture only a fraction of the global B. bifidum pangenome. Methods: In this study, we reconstructed the first comprehensive pangenome of B. bifidum using 1,351 high-quality genomes, including metagenome-assembled genomes. This dataset was used for in silico comparative genomics analyses to identify species-specific genetic and functional features. In vitro transcriptomics analyses were further performed to validate and functionally characterize selected species-specific traits. Results: Comparative genomic analysis with other human-associated bifidobacteria species identified 667 B. bifidum-specific clusters of orthologous genes mostly involved in carbohydrate utilization, osmotic regulation, and host interaction. Notably, B. bifidum displays the most extensive enzymatic repertoire for host-glycan degradation, dedicating 43% of its conserved glycoside hydrolases to these substrates. We identified significant gain-of-function events, including two unique phosphotransferase systems (PTS) for disaccharide uptake. Transcriptomic profiling corroborated the functional relevance of these PTS clusters, which were significantly up-regulated during growth on human milk oligosaccharides, mucin, and N-acetylglucosamine. While the species exhibits high genomic stability, a localized divergence (average nucleotide identitiy, ANI < 98.5%) was identified in rural, non-Westernized populations, reflecting niche-specific adaptations. Conclusion: The identified genomic framework highlighted a distinct evolutionary path of B. bifidum, placing this taxon as a metabolic cornerstone in the neonatal gut via extensive metabolic specialization toward glycan hosts.

Emanuele Selleri, G. Longhi, C. Tarracchini et al. · 0 citations
Open access Jul 2026

Refining Salinivibrio pangenome dynamics and biotechnological potential through comparative analysis

Abstract Current understanding of genomic diversity within the halophilic genus Salinivibrio relies predominantly on draft genomes, with only seven complete genomes among the 62 publicly available. Previous pangenome analysis suggested a closed genomic structure while concluding that Salinivibrio lacks polyhydroxyalkanoate (PHA) degradation capacity despite possessing biosynthesis genes. Here, we present eight complete Salinivibrio genomes from Pearse Lakes (Rottnest Island, Western Australia) generated using Oxford Nanopore long-read sequencing, alongside re-analysis of 38 high-quality public genomes (≥90% completeness and ≤5% contamination cut-off). Pangenome analysis revealed a more open structure than previously reported, with a core genome comprising 25% of total gene clusters and an accessory genome accounting for 71%. Panstripe analysis demonstrated significant temporal signal in gene gain and loss events associated with phylogenetic branch length (core: P=1.72×10⁻⁴; tip: P=2.64×10⁻¹⁴). All 46 genomes contained complete PHA biosynthesis operons (phaB-phaA-phaP-phaC) with high sequence conservation under strong purifying selection (Z=30.30, P<0.001). In a genome that readily gains and loses genes, this conservation indicates that PHA synthesis is a maintained pathway, which is difficult to reconcile with a previous report that Salinivibrio lacks PHA degradation capacity. We therefore searched the genomes by Hidden Markov Model-based homology rather than standard annotation and identified seven putative depolymerases that form a single accessory cluster in 15% of strains, all previously annotated as 3-oxoadipate enol-lactonase-2. These candidates retained all catalytic residues characteristic of active depolymerases but are divergent from reference PHA depolymerases which could explain why annotation missed them. They remain putative and require biochemical confirmation. Both the expanded pangenome and these candidates emerged from standardized homology-based re-analysis, showing that annotation-dependent approaches can overlook genomic diversity and divergent enzyme families in non-model organisms. Together, these results establish Salinivibrio as a genomically dynamic genus with potential for halophilic bioplastic production.

Crystal E. Young, Harrison O'Sullivan, Hussain Alattas et al. · 0 citations