Skip to content
Review

Scalable Extraction and Normalization of Biomedical Knowledge from Research Literature

2026 · AHFE International · 0 citations

TL;DR

A knowledge graph generation pipeline that extracts subject-predicate-object triplets representing scientific claims from research articles and implements a multi-stage consolidation process focused on normalizing entities and relations, offering a robust framework for reducing ambiguity and improving the interoperability of the resulting knowledge graphs.

Abstract

The growth of biomedical literature poses significant challenges for researchers conducting systematic and scoping reviews. In fields such as the use of digital biomarkers for the treatment of heart failure and cardiovascular disease, manually screening thousands of papers is time-consuming and not scalable. To address this problem, we developed AI-assisted tools for large scale analysis and structured knowledge extraction from biomedical research articles. Our approach emphasizes cross-graph analytics, facilitating the exploration of relationships among key biomedical concepts, including digital biomarkers and digital health technologies. By automating the extraction and structuring of knowledge from unstructured text, our system aims to accelerate evidence of synthesis and support more comprehensive and up-to-date reviews in rapidly evolving biomedical domains.We developed a knowledge graph generation pipeline that extracts subject-predicate-object triplets representing scientific claims from research articles. To address redundancy caused by linguistic variation across documents, we implemented a multi-stage consolidation process focused on normalizing entities and relations. This process begins by validating and filtering extracted triplets, then applies lexical normalization to unify entity representations by removing stop words, resolving variants, and merging acronyms with their full forms. Entity types and relations are similarly standardized to ensure uniformity and clarity. The pipeline is designed to be modular and extensible, allowing for the integration of additional normalization strategies or domain-specific ontologies as needed.To further consolidate equivalent triplets, we leverage embedding-based semantic similarity, enabling the merging of semantically similar entities and relationships even when expressed differently across sources. Additionally, our pipeline utilizes biomedical ontologies such as RxNorm and MeSH to map entities to standardized concept identifiers. This ontology-based normalization ensures that references to the same biomedical concept are unified, regardless of linguistic or spelling differences. We evaluated our approach on a pilot corpus of 150 biomedical research articles, processing over 2,000 extracted triplets. The normalization pipeline reduced the number of unique entity variants by more than 50%, consolidating these into approximately 900 unique, semantically unified relationships. Manual review of a representative sample indicated entity normalization accuracy in the range of 90-95%. In conclusion, the integration of lexical, semantic, and ontology-based normalization offers a robust framework for reducing ambiguity and improving the interoperability of the resulting knowledge graphs. Moreover, this structured and unified representation of knowledge facilitates systematic reviews, meta-analyses, and data-driven decision-making in biomedical science, while enabling advanced querying, trend analysis, and the identification of novel associations between biomedical concepts

View source

Similar papers

Optimizing large language model prompts for biomedical knowledge discovery

This work presents a scalable, reproducible framework for evaluating, optimizing, and interpreting LLMs for biomedical knowledge extraction, with a focus on gene–gene regulatory relation prediction, pathway component recognition, multimodal pathway figure understanding, and automated prompt optimization.

Muhammad Azam · 0 citations
Open access Aug 2026

A unified framework and benchmark for generalizable biomedical knowledge extraction and applications with large language models

Results demonstrate that InfoFlowEX equips LLMs with robust adaptability, achieving consistent gains over baselines with minimal task-specific customization, highlighting InfoFlowEX for real-world biomedical applications.

Wuyang Lan, Siqi Zhang, Wenzheng Wang et al. · 0 citations
Jul 2026

Benchmarking Resource-Efficient LLMs for Research Topic Ontology Generation in the Biomedical Field

Assessment of five small, open-source LLMs in identifying semantic relationships between biomedical concepts confirms that direct fine-tuning effectively exceeds the reasoning bottlenecks of smaller LLMs, providing an accurate, automated methodology for the construction and evolution of specialised biomedical ontologies.

Tanay Aggarwal, Angelo Salatino, Francesco Osborne et al. · 0 citations
Aug 2026

CGX: OCR-enhanced knowledge graph retrieval for explainable heart failure analysis

Initial experiments on heart-failure-focused clinical question answering show that CGX improves evidence retrieval quality and perceived answer reliability over conventional retrieval methods, while reducing total graph construction time by 69.7% under the same input corpus and hardware setting.

Dat Nguyen, Anh N. Le, Binh T. D. Trinh et al. · 0 citations
Review Aug 2026

From text to insight: a systematic literature review of keyword and keyphrase extraction techniques for healthcare text mining

This review provides researchers and practitioners with a structured framework for method selection based on their specific constraints and identifies six prioritized research directions for future investigation, identifying critical research gaps including the preservation of multi-word clinical concepts, scarce evaluation in real-world clinical workflows, and persistent hallucination risks in model-based extraction.

Mouhamed Gaith Ayadi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.