KRAKEN (Knowledge Research & Analysis Kit for Evidence Networks) addresses this gap by integrating existing graphs with specialized sources such as RefMet, LIPID MAPS, NIH Common Data Elements, Polygenic Score Catalog, and derived wellness measures including biological age and biological BMI.
Abstract
Existing general-purpose biomedical knowledge graphs tend to focus on disease mechanisms and drug repurposing, leaving multiomic and wellness-relevant content underrepresented. KRAKEN (Knowledge Research & Analysis Kit for Evidence Networks) addresses this gap by integrating existing graphs (including Translator KG Open, RTX-KG2, and ROBOKOP) with specialized sources such as RefMet, LIPID MAPS, NIH Common Data Elements, Polygenic Score Catalog, and derived wellness measures including biological age and biological BMI. The resulting graph spans ∼15M nodes and ∼113M edges across 62 entity types. KRAKEN adopts the Biolink Model as its semantic layer, ensuring compatibility with standardized resources emerging from the NIH NCATS Biomedical Data Translator program. A lightweight, modular build system rebuilds the full graph (including entity resolution), with peak memory consumption <48 GB, and supports flexible inclusion or exclusion of sources, allowing the user to scope the graph to a domain of interest. Built-in analytical tools include multi-hop reasoning, subgraph extraction, text, vector and hybrid entity search, and enrichment analyses, all accessible through an interactive web interface, a REST API, and a Model Context Protocol server, the last enabling direct consumption by agentic and LLM-based systems. KRAKEN is freely available at https://app.krakenkg.com. GRAPHICAL ABSTRACT
Identifying causal connections between existing drugs and mechanistic profiles of diseases is a foundational step for effective drug repurposing. Although knowledge graphs (KGs) are highly suited for consolidating biomedical databases and tracking these connections, a single biomedical KG is constrained by its ingestion pipeline and knowledge sources. While different biomedical KGs could be complementary if combined, efforts to combine them into a unified and more comprehensive KG are hindered by lack of interoperability and poor provenance. To address those issues, we present EC-KG, a Biolink Model-compatible KG for computational drug repurposing. EC-KG is an interoperable, provenance-first KG which integrates RTX-KG2, ROBOKOP, and PrimeKG at the network-level, encapsulating over 7 million nodes and 81 million edges from 95 primary data sources. EC-KG has improved coverage of core biomedical entities such as drugs, targets, and diseases relevant to drug repurposing vs source graphs, and captures complex biomedical mechanisms within its topology. We demonstrate that the network unification in EC-KG leads to emergence of novel, mechanistically relevant pathways which are disconnected in the underlying constituent networks and show its applications in method development, benchmarking and predictive drug repurposing applications. EC-KG has already been successfully used in drug repurposing research to surface Botulinum Toxin A as a candidate to treat Major Depressive Disorder, as well as to validate repurposing of Lenalidomide and Dexamethasone for a subgroup of patients with Rosai-Dorfman Disease.
Piotr Kaniewski, E. K. Carter, Daniel J. Rhodes et al.· bioRxiv· 0 citations
Initial experiments on heart-failure-focused clinical question answering show that CGX improves evidence retrieval quality and perceived answer reliability over conventional retrieval methods, while reducing total graph construction time by 69.7% under the same input corpus and hardware setting.
Dat Nguyen, Anh N. Le, Binh T. D. Trinh et al.· Journal of Biomedical Inform...· 0 citations
Chronic diseases impose a significant burden on global health. Knowledge graphs (KGs), which integrate and represent multisource information as interconnected networks, provide a promising approach for enabling personalized and dynamic health management. The aim of this scoping review is to systematically map the current landscape of KG applications in chronic disease health management.
In accordance with the Arksey and O’Malley framework and the PRISMA-ScR guidelines, a systematic search was conducted in eight databases, namely, the Wanfang Database, CNKI, VIP, SinoMed, PubMed, Embase, Web of Science, and CINAHL, from the establishment of the databases to August 2025. Two researchers independently screened the literature and extracted data on the basis of the inclusion and exclusion criteria.
A total of 15 studies published between 2018 and 2025 were included. In these studies, the KG construction process generally involved five stages: data collection, information extraction, knowledge fusion, graph construction, and visualization. KGs were applied across a range of chronic diseases, including metabolic, cardiovascular, and respiratory diseases, as well as cancers. In terms of application scenarios, self-health management was the most common (9 studies, 60.0%), followed by clinical decision support (5 studies, 33.3%), intelligent question answering (4 studies, 26.7%), and health monitoring/early warning (1 study, 6.7%). With respect to system maturity, 12 studies (80.0%) were classified as prototype systems, 3 studies (20.0%) were classified as pilot implementations, and none were classified as clinically deployed systems. The reported outcomes were mainly technical or feasibility oriented, and 7 studies (46.7%) did not report any explicit evaluation metrics.
Knowledge graphs show promise for integrating heterogeneous health data in chronic disease management. However, research has focused predominantly on system development and proof-of-concept validation, with limited high-quality evidence regarding clinical effectiveness. To support clinically meaningful implementation, in future work, rigorous real-world evaluation should be prioritized and challenges related to data quality, interoperability, and model interpretability should be addressed.
Ziting Xu, Chunhao Dai, Haoyuan Li et al.· BMC Medical Informatics and...· 0 citations
This work introduces MolBioKG, a two-layer system that grounds unseen molecules in biomedical evidence via multi-resolution structural anchoring and outperforms strong baselines across in-graph link recovery, complex multi-hop reasoning, and out-of-graph generalization.
Yiming Zhang, Hikaru Shindo, Shuan Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.