Skip to content
Open access

Language-Model-Based Architecture for Automatic Concept Placement in Ontologies

Aug 2026 · Information · 0 citations · 19 references

TL;DR

This paper addresses the placement of concepts that are absent from the target ontology—the out-of-knowledge-base setting—in which a textual mention must be assigned one or more insertion positions in the subsumption hierarchy rather than linked to an existing node.

Abstract

Integrating newly emerging terms into existing ontologies is a recurring maintenance problem in knowledge engineering, particularly in biomedical domains where terminology evolves faster than manual curation can accommodate. This paper addresses the placement of concepts that are absent from the target ontology—the out-of-knowledge-base setting—in which a textual mention must be assigned one or more insertion positions in the subsumption hierarchy rather than linked to an existing node. We propose a three-stage framework that extends the conventional retrieve-then-select paradigm with an intermediate stage of edge generation and enrichment, which expands the candidate set by traversing the local structure of the ontology. Stage 1 retrieves candidate edges using a fine-tuned bi-encoder trained with a max-margin objective; Stage 2 constructs and structurally enriches candidate edges; Stage 3 selects among them using either a fine-tuned cross-encoder or a large language model under explainable instruction tuning. We evaluate on two datasets derived from SNOMED CT, MM-S14-Disease and MM-S14-CPP, under a strict out-of-knowledge-base protocol. Fine-tuned pre-trained language models outperform zero-shot and instruction-tuned large language models on ranking accuracy, while the instruction-tuned configuration produces expert-auditable justifications at a modest cost in accuracy. On MM-S14-Disease, the strongest configuration places a correct insertion edge among the ten highest-ranked candidates for 38.7% of test mentions and recovers the complete gold edge set for 16.4%, against 26.1% and 9.2% for retrieval alone. The framework is positioned as decision support for ontology curators rather than as an autonomous ontology generator.

Read PDF

Similar papers

Jul 2026

An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents

A production extraction layer that converts a live document stream into a validated knowledge graph aligned to a formal ontology, and improved search recall from roughly 70 to 95 percent with no false merges, and corrected seven classes of silent quality defect.

Vaibhav Dangaich, Kevin Lewis, Kundeshwar Pundalik · 0 citations
Open access Jul 2026

Integrating Heterogeneous Knowledge for Enhanced Recommendation with Large Language Models

The proposed REKALM, a comprehensive integration framework for enhancing LLM-based recommenders through knowledge integration, demonstrates that augmenting LLMs with lexicalized, domain-specific knowledge is an effective system-level strategy for advancing the next generation of recommender systems.

Alessandro Petruzzelli, C. Musto, Marco De Gemmis et al. · 0 citations
Jul 2026

Benchmarking Resource-Efficient LLMs for Research Topic Ontology Generation in the Biomedical Field

Assessment of five small, open-source LLMs in identifying semantic relationships between biomedical concepts confirms that direct fine-tuning effectively exceeds the reasoning bottlenecks of smaller LLMs, providing an accurate, automated methodology for the construction and evolution of specialised biomedical ontologies.

Tanay Aggarwal, Angelo Salatino, Francesco Osborne et al. · 0 citations
Jul 2026

LLM-Assisted Ontology Engineering and Construction of a French Legal Knowledge Graph

A two-stage LLM-assisted workflow for French maintenance regulations is presented: ontology engineering from a SEMLEG-based core ontology, followed by construction of an ontology-grounded French legal knowledge graph.

Génesis Montenegro, M. Billami, Catherine Faron et al. · 0 citations
#artificial intelligence Preprint Aug 2026

When Does Bigger Help? A Controlled Study of LLM Scale for Ontology Learning

The effect of Large Language Model (LLM) scale on ontology learning (OL) performance remains insufficiently characterized. We present a controlled evaluation of 13 models spanning dense and Mixture-of-Experts variants from the Qwen3.5 and Qwen3.6 lineages, together with proprietary GPT release variants, using the OntoLearner retrieval-augmented generation pipeline. All models are evaluated with the same embedding model, retrieval configuration, prompt templates, decoding settings, datasets, and metrics on term typing, taxonomy discovery, and non-taxonomic relationship extraction across four biomedical and materials science and engineering ontologies. Within the dense Qwen3.5 lineage, increasing parameter count primarily improves precision rather than recall, with the largest gains occurring between 9B and 27B parameters. However, the effect of scale is neither monotonic nor uniform across tasks and domains. Dense 27B models outperform substantially larger sparse models on term typing, whereas larger Mixture-of-Experts models achieve the strongest open-weight results on taxonomy discovery. Non-taxonomic relationship extraction remains difficult across model scales, particularly for the Materials Data Science ontology. Performance differences across matched Qwen variants and proprietary GPT releases further indicate that architecture and model lineage can outweigh nominal parameter count. These findings show that model size alone is an insufficient selection criterion for OL and provide empirical guidance for reproducible LLM-assisted ontology engineering.

Hamed Babaei Giglou, S. Auer, Jennifer D’Souza · 0 citations
Open access Aug 2026

Optimizing sample selection for large language model-based entity matching using AssistEM

AssistEM, a framework for efficient LLM adaptation to EM via principled data selection, demonstrates that selective fine-tuning not only accelerates adaptation but also improves training efficiency (requiring fewer GPU hours), enabling open-source LLMs to rival–and in some cases outperform–closed-source models.

John Bosco Mugeni, Steven J. Lynden, Toshiyuki Amagasa et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.