Assessment of five small, open-source LLMs in identifying semantic relationships between biomedical concepts confirms that direct fine-tuning effectively exceeds the reasoning bottlenecks of smaller LLMs, providing an accurate, automated methodology for the construction and evolution of specialised biomedical ontologies.
Abstract
Knowledge Organization Systems like Ontologies and taxonomies are fundamental for structuring scientific knowledge, yet their manual curation presents a persistent bottleneck in knowledge management. While Large Language Models (LLMs) offer a scalable mechanism for automated ontology generation, their capacity to classify complex, domain-specific semantics requires systematic evaluation. In this paper, we assess the performance of five small, open-source LLMs (up to 9 billion parameters) in identifying semantic relationships between biomedical concepts. To support this evaluation, we introduce MeSH-Rel-4K, a dataset comprising 4K semantic relationships extracted from the Medical Subject Headings (MeSH). We analyse three adaptation strategies: standard prompting, Chain-of-Thought prompting, and fine-tuning. While parameter-constrained models traditionally struggle with the nuances of in-context logic, our results reveal that targeted fine-tuning increases the average F1-score by 34.1 percentage points. These results confirm that direct fine-tuning effectively exceeds the reasoning bottlenecks of smaller LLMs, providing an accurate, automated methodology for the construction and evolution of specialised biomedical ontologies.
The effect of Large Language Model (LLM) scale on ontology learning (OL) performance remains insufficiently characterized. We present a controlled evaluation of 13 models spanning dense and Mixture-of-Experts variants from the Qwen3.5 and Qwen3.6 lineages, together with proprietary GPT release variants, using the OntoLearner retrieval-augmented generation pipeline. All models are evaluated with the same embedding model, retrieval configuration, prompt templates, decoding settings, datasets, and metrics on term typing, taxonomy discovery, and non-taxonomic relationship extraction across four biomedical and materials science and engineering ontologies. Within the dense Qwen3.5 lineage, increasing parameter count primarily improves precision rather than recall, with the largest gains occurring between 9B and 27B parameters. However, the effect of scale is neither monotonic nor uniform across tasks and domains. Dense 27B models outperform substantially larger sparse models on term typing, whereas larger Mixture-of-Experts models achieve the strongest open-weight results on taxonomy discovery. Non-taxonomic relationship extraction remains difficult across model scales, particularly for the Materials Data Science ontology. Performance differences across matched Qwen variants and proprietary GPT releases further indicate that architecture and model lineage can outweigh nominal parameter count. These findings show that model size alone is an insufficient selection criterion for OL and provide empirical guidance for reproducible LLM-assisted ontology engineering.
Hamed Babaei Giglou, S. Auer, Jennifer D’Souza· 0 citations
LLMBDC (Large Language Model for Biological Domains Oriented Clustering of Gene Ontology) provides a scalable, reproducible, and interpretable route to context-aware, system-level interpretation of GO enrichment results while preserving biological specificity.
Results demonstrate that InfoFlowEX equips LLMs with robust adaptability, achieving consistent gains over baselines with minimal task-specific customization, highlighting InfoFlowEX for real-world biomedical applications.
Wuyang Lan, Siqi Zhang, Wenzheng Wang et al.· Cell Reports Medicine· 0 citations
Biomedical ontology normalization maps free-text expressions to standardized concepts, enabling consistent integration and analysis of biomedical data. This task remains challenging because lexical variation and subtle distinctions among hierarchically related concepts can obscure concept boundaries. We present OntologyAligner, a three-stage framework that combines ontology-aligned retrieval, large language model candidate reranking, and selective hierarchy-guided refinement. We also construct PhenoNormBench, a unified benchmark comprising 13,390 samples from seven Human Phenotype Ontology datasets. OntologyAligner achieved state-of-the-art performance on HPO normalization, with 88.78% Macro Top-1 Accuracy and 86.75% Micro Top-1 Accuracy, exceeding the strongest baseline by 4.85 and 5.07 percentage points, respectively. Ablation analyses showed complementary contributions from all three stages, and sensitivity analyses demonstrated stability across candidate-set sizes and model backbones. Applications to MONDO, MEDIC, and NCBITaxon further established portability to other ontologies. OntologyAligner offers a generalizable framework for accurate mapping of biomedical text to structured ontology concepts. PhenoNormBench and the code are publicly available at https://github.com/zhelishisongjie/OntologyAligner.
Jie Song, Zhichuan Xu, Ziyu Lu et al.· 0 citations
Summary Diagnosing rare diseases remains a major challenge due to limited clinical knowledge and the frequent absence of diagnostic criteria. We present a digital framework that leverages large language models and biomedical text embeddings to bridge this gap. By mapping Human Phenotype Ontology terms to a shared vector space with millions of PubMed abstracts and full-text articles, our method enables phenotype-driven semantic search and ranks literature relevant to patient symptoms, even without explicit disease mentions. Validated on OMIM-derived benchmarks and applied to RASopathies, including NF1, Noonan, and Costello syndromes, our approach retrieved expected findings, supporting differential diagnosis and research. The framework is implemented in an open-source Python package, py-semtools, and it can be integrated into clinical decision support systems or adapted to other ontologies and corpora. This work demonstrates how AI-driven informatics can enhance rare disease diagnosis and exemplifies the role of digital tools in transforming precision medicine and healthcare delivery.
Jesús Pérez-García, Federico García-Criado, F. Pazos et al.· iScience· 0 citations
Ontology enrichment is a critical but labor-intensive step in semantic knowledge representation. To address this challenge, we propose OntoCodex, a multi-agent framework that integrates large language models (LLMs), ontologies, curated knowledge sources, and standard vocabularies to support semi-automated ontology enrichment with formal OWL-based integration and a feedback loop. OntoCodex consists of five coordinated agents for ontology parsing, task decision-making, knowledge retrieval, terminology normalization, and automated script generation. We evaluated OntoCodex using a ChatGPT-4o–powered implementation to enrich concepts across five chronic diseases, including stroke, chronic obstructive pulmonary disease, atrial fibrillation, osteoporosis, and Parkinson’s disease. Compared with baseline ChatGPT-4o, OntoCodex improved concept extraction across most domains, achieving higher precision, recall, and F1 scores, including perfect performance in laboratory test extraction, and demonstrated greater accuracy in standardized terminology mapping, particularly for medications, while showing lower performance in laboratory test mapping. Automatically generated Python scripts successfully enriched the MCC-CDO with new concepts and annotations without errors. These results demonstrate that OntoCodex substantially improves ontology enrichment and has strong potential to accelerate clinical and translational research.
Jingna Feng, Yue Yu, Aaron Dong et al.· npj Health Systems· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.