The findings demonstrate that Gene Ontology can be effectively leveraged for semantics-aware gene selection in clinical machine learning and offers a scalable strategy to uncover new, testable biological hypotheses–revealing gene functions that might otherwise remain hidden when examining only broadly selected gene combinations.
Abstract
The Gene Ontology (GO) is a public resource that describes gene functions and characteristics through a structured vocabulary of standardised terms. It currently contains annotations for over 1.5 million gene products, each linked to one or more GO terms. In this study, we propose integrating this GO-based semantic structure into machine learning systems for medical diagnostics. This approach serves a dual purpose: first, to prioritise genes that are semantically relevant to a given clinical task, thereby refining model input; and second, to enable the analysis of biologically predefined gene sets, which may reveal novel mechanisms underlying disease.
Evaluated across 16 benchmark data sets spanning diverse medical domains, our GO term-informed gene selection method generally outperformed models trained on full gene sets. Further analysis of individual GO terms not only enhanced classification performance but also identified high-performing, task-specific gene subsets that were overlooked during initial gene selection.
Our findings demonstrate that Gene Ontology can be effectively leveraged for semantics-aware gene selection in clinical machine learning. Moreover, systematically evaluating individual GO terms offers a scalable strategy to uncover new, testable biological hypotheses–revealing gene functions that might otherwise remain hidden when examining only broadly selected gene combinations.
LLMBDC (Large Language Model for Biological Domains Oriented Clustering of Gene Ontology) provides a scalable, reproducible, and interpretable route to context-aware, system-level interpretation of GO enrichment results while preserving biological specificity.
An embedding-based statistical framework is developed that transforms gene and function descriptions into vector representations, enabling statistical testing of gene-gene and gene-function relationships and quantitative prioritization of de novo functional hypotheses inferred by LLMs.
Summary Diagnosing rare diseases remains a major challenge due to limited clinical knowledge and the frequent absence of diagnostic criteria. We present a digital framework that leverages large language models and biomedical text embeddings to bridge this gap. By mapping Human Phenotype Ontology terms to a shared vector space with millions of PubMed abstracts and full-text articles, our method enables phenotype-driven semantic search and ranks literature relevant to patient symptoms, even without explicit disease mentions. Validated on OMIM-derived benchmarks and applied to RASopathies, including NF1, Noonan, and Costello syndromes, our approach retrieved expected findings, supporting differential diagnosis and research. The framework is implemented in an open-source Python package, py-semtools, and it can be integrated into clinical decision support systems or adapted to other ontologies and corpora. This work demonstrates how AI-driven informatics can enhance rare disease diagnosis and exemplifies the role of digital tools in transforming precision medicine and healthcare delivery.
Jesús Pérez-García, Federico García-Criado, F. Pazos et al.· iScience· 0 citations
Assessment of five small, open-source LLMs in identifying semantic relationships between biomedical concepts confirms that direct fine-tuning effectively exceeds the reasoning bottlenecks of smaller LLMs, providing an accurate, automated methodology for the construction and evolution of specialised biomedical ontologies.
Tanay Aggarwal, Angelo Salatino, Francesco Osborne et al.· arXiv.org· 0 citations
Benchmarks for rare-disease gene prioritisation are assembled from published clinical cases. Those cases often come from the same publications used to build knowledge-base (“curated”) tools, so a curated tool can be scored on its own source literature. This resource makes that circularity measurable. We release a stratified benchmark of 1,047 rare-disease cases from the GA4GH Phenopacket Store v0.1.26. Each case pairs a Human Phenotype Ontology profile with a 50-gene candidate list (one causal gene, 49 distractors) and the causal-gene label, sampled across four operational MONDO-derived disease strata and issued in two case-paired variants: random distractors, and phenotype-similar distractors selected by HPO Resnik similarity. Two case-level metadata layers support fairer evaluation: a per-case flag recording whether a case’s source publication is cited in the HPO disease-annotation file, defining an overlap-absent subset (n = 282), and publication-recency strata. We also specify a deterministic, version-pinned recipe for a hybrid dense-plus-sparse retrieval index over ∼ 2.25 million PMC Open Access articles (52,777,395 chunks). The resource reports no tool comparisons.
This work presents a scalable, reproducible framework for evaluating, optimizing, and interpreting LLMs for biomedical knowledge extraction, with a focus on gene–gene regulatory relation prediction, pathway component recognition, multimodal pathway figure understanding, and automated prompt optimization.
Muhammad Azam· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.