LLMBDC (Large Language Model for Biological Domains Oriented Clustering of Gene Ontology) provides a scalable, reproducible, and interpretable route to context-aware, system-level interpretation of GO enrichment results while preserving biological specificity.
Abstract
Gene Ontology (GO) enrichment analysis is a foundational tool for translating large-scale genomic data into biological insights, but typically yields hundreds of redundant terms that obscure overarching themes. Existing summarization tools rely on fixed similarity metrics (REVIGO, GOSemSim, clusterProfiler::simplify()), gene-overlap measures (Metascape), or static hierarchy mappings (GO-slim), and therefore cannot incorporate biological context. Manual curation provides context-aware grouping but is subjective and labor-intensive. A scalable, context-aware framework is needed to cluster GO terms into interpretable higher-order biological domains. Here we present LLMBDC (Large Language Model for Biological Domains Oriented Clustering of Gene Ontology), a training-free framework that leverages zero-shot semantic reasoning of LLMs with confidence scoring to cluster GO terms into BioDomains using only ontology information at inference time. Benchmarked across Alzheimer's disease (AD) and Fragile X syndrome (FXS) against six baseline methods including SapBERT, LLMBDC achieved substantially higher precision, recall, and clustering performance. Against ground-truth annotations, LLMBDC improved ARI from 9.7% to 73.3% (AD) and from 15.7% to 66.6% (FXS) over REVIGO, with corresponding NMI gains from 59.9% to 73.4% (AD) and 66.0% to 79.5% (FXS). A Cauchy combination test further confirmed that aggregated BioDomains retained statistically significant functional signals. LLMBDC provides a scalable, reproducible, and interpretable route to context-aware, system-level interpretation of GO enrichment results while preserving biological specificity.
The findings demonstrate that Gene Ontology can be effectively leveraged for semantics-aware gene selection in clinical machine learning and offers a scalable strategy to uncover new, testable biological hypotheses–revealing gene functions that might otherwise remain hidden when examining only broadly selected gene combinations.
Peter Eckhardt-Bellmann, N. Taha, Silke D. Werle et al.· BMC Bioinformatics· 0 citations
Assessment of five small, open-source LLMs in identifying semantic relationships between biomedical concepts confirms that direct fine-tuning effectively exceeds the reasoning bottlenecks of smaller LLMs, providing an accurate, automated methodology for the construction and evolution of specialised biomedical ontologies.
Tanay Aggarwal, Angelo Salatino, Francesco Osborne et al.· arXiv.org· 0 citations
This study explores whether human-written descriptions in Reactome can be used to infer the experts'defined global hierarchical structure and indicates that the global hierarchical structure of pathways can be inferred by experts textual metadata.
Susanna Bravi, R. De Luca, R. Sicilia et al.· 0 citations
The effect of Large Language Model (LLM) scale on ontology learning (OL) performance remains insufficiently characterized. We present a controlled evaluation of 13 models spanning dense and Mixture-of-Experts variants from the Qwen3.5 and Qwen3.6 lineages, together with proprietary GPT release variants, using the OntoLearner retrieval-augmented generation pipeline. All models are evaluated with the same embedding model, retrieval configuration, prompt templates, decoding settings, datasets, and metrics on term typing, taxonomy discovery, and non-taxonomic relationship extraction across four biomedical and materials science and engineering ontologies. Within the dense Qwen3.5 lineage, increasing parameter count primarily improves precision rather than recall, with the largest gains occurring between 9B and 27B parameters. However, the effect of scale is neither monotonic nor uniform across tasks and domains. Dense 27B models outperform substantially larger sparse models on term typing, whereas larger Mixture-of-Experts models achieve the strongest open-weight results on taxonomy discovery. Non-taxonomic relationship extraction remains difficult across model scales, particularly for the Materials Data Science ontology. Performance differences across matched Qwen variants and proprietary GPT releases further indicate that architecture and model lineage can outweigh nominal parameter count. These findings show that model size alone is an insufficient selection criterion for OL and provide empirical guidance for reproducible LLM-assisted ontology engineering.
Hamed Babaei Giglou, S. Auer, Jennifer D’Souza· 0 citations
Ontology enrichment is a critical but labor-intensive step in semantic knowledge representation. To address this challenge, we propose OntoCodex, a multi-agent framework that integrates large language models (LLMs), ontologies, curated knowledge sources, and standard vocabularies to support semi-automated ontology enrichment with formal OWL-based integration and a feedback loop. OntoCodex consists of five coordinated agents for ontology parsing, task decision-making, knowledge retrieval, terminology normalization, and automated script generation. We evaluated OntoCodex using a ChatGPT-4o–powered implementation to enrich concepts across five chronic diseases, including stroke, chronic obstructive pulmonary disease, atrial fibrillation, osteoporosis, and Parkinson’s disease. Compared with baseline ChatGPT-4o, OntoCodex improved concept extraction across most domains, achieving higher precision, recall, and F1 scores, including perfect performance in laboratory test extraction, and demonstrated greater accuracy in standardized terminology mapping, particularly for medications, while showing lower performance in laboratory test mapping. Automatically generated Python scripts successfully enriched the MCC-CDO with new concepts and annotations without errors. These results demonstrate that OntoCodex substantially improves ontology enrichment and has strong potential to accelerate clinical and translational research.
Jingna Feng, Yue Yu, Aaron Dong et al.· npj Health Systems· 0 citations
GEOMeta provides a scalable resource and reproducible framework for metadata curation in the Gene Expression Omnibus, and benchmarked transcriptome representation models for predicting sex, age, tissue and disease from transcriptome embeddings.
Xiaodan Zhang, S. Paithankar, Jing Pu et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.