DrugBank in RDF: Vector Embeddings
Abstract
This work addresses the need for efficient and reusable vector embeddings (VEs) for large-scale RDF Knowledge Graphs, focusing on the widely used DrugBank dataset. The study aims to reduce the computational burden and environmental impact associated with repeatedly generating embeddings for downstream tasks such as link prediction, clustering, and recommendation. To this end, embeddings are generated from the DrugBank Knowledge Graph—comprising over 3.6 million triples, more than 1.5 million entities, and 95 relations—using 25 models implemented in the PyKEEN framework. The dataset is processed into training, validation, and test splits, and embeddings of fixed dimensionality are produced for both entities and relations. The resulting representations are released in multiple formats, including full JSON files, class-partitioned subsets, and an efficient Parquet-based structure, to support scalable querying. Experimental evaluation demonstrates substantial improvements in access time and memory usage when using structured formats, particularly Parquet. Additionally, the study quantifies the carbon footprint of embedding generation, showing that distributing precomputed embeddings can reduce energy consumption by up to 99.37% compared to on-demand recomputation. Overall, the work provides a comprehensive, reusable resource that facilitates research while promoting computational efficiency and environmental sustainability.