Aug 2026· Microorganisms· Vol 14· 0 citations· 34 references
Medicine
TL;DR
Context-Aware Annotation Completion (CAAC), a framework integrating ESM2 embeddings, genomic-neighborhood features, three-class classification, confidence-tiered neighbor voting, and Enzyme Commission (EC)-to-KEGG Orthology (KO) mapping, extends the enzyme-level interpretation of under-annotated silage metagenomes, while the inferred assignments remain computational predictions requiring experimental validation.
Abstract
Functional annotation gaps limit the interpretation of carbohydrate metabolism in silage microbiomes. We developed Context-Aware Annotation Completion (CAAC), a framework integrating ESM2 embeddings, genomic-neighborhood features, three-class classification, confidence-tiered neighbor voting, and Enzyme Commission (EC)-to-KEGG Orthology (KO) mapping. CAAC was applied to 21 metagenomes from uninoculated and Lacticaseibacillus paracasei-inoculated silages sampled before ensiling and at 7 and 90 days. Five-fold cross-validation yielded an F1-macro of 84.64% for negative, positive, and hard-sequence classification. Among 800,000 selected annotation-poor sequences, 545,671 Tier 1 or Tier 2 predictions passed the annotation-validity and EC-to-KO mapping criteria, of which 524,814 were eligible for sample-level annotation supplementation. After silage-focused filtering and KO–EC summarization, these predictions yielded 102 KO–EC features repeatedly detected across the silage metagenomes and increased coverage in 25 of 47 carbohydrate-metabolism pathways, mainly by recovering enzyme-level components related to starch and sucrose, cellulose and cellobiose, xylan and hemicellulose, and pectin and glucuronate metabolism. Taxon-linked analyses further revealed treatment- and stage-associated patterns in the taxonomic sources of the supplemented annotations. A database-derived temporal benchmark using the July 2025 CAZy release showed 94.94% Tier 1 family-level annotation-transfer consistency. CAAC extends the enzyme-level interpretation of under-annotated silage metagenomes, while the inferred assignments remain computational predictions requiring experimental validation.
Metagenomics enables the recovery of metagenome-assembled genomes (MAGs), providing access to the metabolic potential of uncultured microbial communities that drive ecosystem function and biogeochemical cycles. However, as MAGs datasets increase in size and complexity, comparing functional repertoires and identifying ecologically meaningful traits across experimental gradients becomes increasingly difficult. Here, we present rbims, a modular R package for integrative functional profiling of MAGs and metagenomic datasets. rbims supports annotations from KEGG, dbCAN, InterProScan, MEROPS, and PICRUSt2, and enables the calculation of gene presence/absence, raw abundance, and pathway coverage, as well as metadata-informed comparative analyses and publication-ready visualizations. Beyond descriptive profiling, rbims implements an exploratory discriminant framework that combines compositional differential analysis (ALDEx2) with random forest–based feature ranking to prioritize candidate metabolic traits associated with environmental factors. Importantly, it extends gene-level analysis to pathway-level directional bias testing, allowing users to evaluate whether the majority of genes within a metabolic route are consistently enriched toward a given condition. We applied rbims to 42 MAGs recovered from a hydrocarbon enrichment experiment in the North Atlantic Ocean. The workflow identified widespread hexadecane and phenanthrene degradation potential, detected enriched oxidoreductase-related protein families, and revealed a strong pathway-level directional bias toward deep-water MAGs for phenanthrene, naphthalene, and hexadecane degradation pathways. By integrating annotation parsing, quantitative trait analysis, statistical discrimination, and visualization in a reproducible framework, rbims provides a user-friendly platform for functional interpretation in genome-resolved metagenomics.
Karla P. López-Martínez, S. Hereira-Pacheco, Diana Hernández-Oaxaca et al.· Frontiers in Bioinformatics· 0 citations
Abstract Motivation Functional characterization of microbiomes often relies on the sequencing of metagenomic DNA extracted from environmental samples, with current approaches using metagenome-assembled genomes (MAGs). Although glycoside hydrolases (GHs) are central to carbon cycling, accurate annotation of GHs in metagenomic datasets remains challenging due to the multidomain architecture of carbohydrate-active enzymes and the prevalence of unassembled short reads due to limitations in the MAG-generation process. Results Here, we present CAZyOGH (CAZymes Open-source GH annotation), a curated reference database for the domain-specific identification of 135 protein domains spanning 99 GH families with well-defined catalytic domain signatures. CAZyOGH focuses on individual GH domains, enabling robust annotation of both assembled and unassembled metagenomic data. We validated CAZyOGH by reanalyzing genomes listed in CAZy db, where predicted GH profiles closely matched reported values. Next, we used CAZyOGH to analyze 12 human gut metagenomes and 12 newly sequenced soil microbiomes to reveal environment-specific GH repertoires. By accurately detecting catalytic domains independent of the genomic context, CAZyOGH improves sensitivity and specificity in short-read metagenomic annotation. This framework provides a scalable and reproducible approach to investigate carbohydrate-active enzymes across ecosystems, advancing our capacity to characterize microbial functional potential in global carbon cycling. Availability and implementation CAZyOGH data is available on figshare (https://figshare.com/projects/CAZyO_GH/267770).
N. Griffin, Alison E Hughes, D. S. Erdody et al.· Bioinformatics Advances· 0 citations
Metagenomic data have notable biological potential, but their functional interpretation is frequently impeded by incomplete protein function annotations. Accurate enzyme annotation is essential for elucidating the metabolic capabilities of microbial communities within metagenomic datasets. To address this challenge, we developed FEDKEA, an enzyme annotation tool leveraging protein language models, and provided a web platform for its use. In addition, we designed a user-friendly, FEDKEA-based metagenomic pipeline, MEnzMap, which encompasses the entire analysis workflow—from raw data quality control to function prediction and downstream analyses. Applying MEnzMap to human gut metagenomic data from the iHMP2 project, we generated a comprehensive enzyme profile landscape for both healthy individuals and patients with inflammatory bowel diseases. These tools provide an efficient method for the functional annotation of microbial dark matter and facilitate the identification of disease-associated enzymes.
Lei Zheng, Bowen Li, Siqi Xu et al.· Science Advances· 0 citations
Microbiomes are information-rich biological systems, yet most computational analyses still reduce communities to cohort-specific abundance tables. Here we introduce MGM2, a multimodal foundation model pretrained on 1,821,291 MicrobeAtlas samples and 225,067 OTUs clustered at 99% sequence similarity. MGM2 couples NTv3-derived microbial sequence embeddings with abundance conditioning and community-semantic alignment to learn transferable sample– and token-level representations. Frozen MGM2 representations outperformed DeepPhylo by 0.06–0.21 macro-AUROC across five temporally held-out MGnify hierarchy levels, with the largest gains for rare and fine-grained labels. In fecal microbiota transplantation, MGM2-XLarge achieved a response ROC AUC of 0.79 and reduced post-transplant Bray-Curtis distance by 15% relative to the recipient baseline. The same representation supported ASV-level trend forecasting across 24 wastewater treatment plants. Sparse autoencoder analysis resolved MGM2-XLarge token states into a 4,096-feature dictionary spanning taxonomic identity, abundance state, ecological context and technical variation. MGM2 therefore provides a sequence-aware and interpretable representation layer for microbiome classification, paired-community prediction, forecasting and feature discovery. Highlights ● MGM2 integrates sequence, abundance and community semantics through pretraining on 1.82 million microbiome samples. ● Frozen MGM2 improved macro-AUROC over DeepPhylo by 0.06–0.21 across five temporally held-out MGnify levels. ● MGM2-XLarge reached a response ROC AUC of 0.79 and reduced post-FMT Bray-Curtis distance by 15%. ● A 4,096-feature sparse autoencoder atlas resolves taxonomic, abundance, ecological and technical signals.
Current bioinformatics approaches for bacterial diagnostic target discovery remain constrained by their reliance on gene annotations and fixed-boundary genome segmentation, which overlook unannotated intergenic regions and introduce sequence-truncation artifacts. Here, we developed an open-source, annotation-independent pan-genomic pipeline featuring an overlapping sliding-window algorithm (500-bp window, 100-bp step) and a three-tier subtractive screening funnel. Using Staphylococcus aureus as a model, the pipeline screened 1,629 genomes against 852 non-S. aureus Staphylococcus genomes and >20,000 background bacterial genomes. Seven highly conserved, unannotated targets (SA-1 to SA-7) were identified, with all seven translated into qPCR primer sets (SAP-1 to SAP-7), among which three (SAP-1 to SAP-3) were further characterized by in vitro experiments. Multi-layer in silico evaluation demonstrated 100% intraspecific sensitivity and zero cross-reactivity against background genomes, including the S. aureus complex. In vitro testing using crude cell lysates confirmed specific amplification of S. aureus DNA without non-target cross-reactivity, establishing a qualitative limit of detection (LOD) of 10
5
CFU/mL. Additional computational validation on draft genomes, raw sequencing reads, near-neighbor species, and a clinical truth set corroborated marker robustness under realistic conditions. This framework successfully circumvents conventional gene-centric limitations, providing a generalizable computational strategy for target discovery across other high-priority bacterial pathogens.
Yuyang Zhou, Jiayi Wang, Junhua Xiao et al.· Journal of Bioinformatics an...· 0 citations