Skip to content

KG-TransomicNet: Semantic–Quantitative Property Graph of PheKnowLator and Multi-Omics

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)
Bioinformatics and Genomic Networks

Abstract

KG-TransomicNet Version 1.1.2 · Software · Open · MIT De Filippis, Giovanni Maria · Rinaldi, Antonio Maria A semantic–quantitative property-graph framework that couples the PheKnowLator biomedical knowledge graph with per-sample multi-omics measurements in ArangoDB, enabling trans-omic analysis and knowledge-based reasoning over TCGA/TARGET cohorts via AQL. What's new in v1.1.2 PanCanAtlas view. A new scripts/pancan_atlas/ package adds the harmonised TCGA PanCanAtlas data (EB++ RNA-seq, batch-adjusted miRNA, iCluster subtype labels) and the pan-cancer GISTIC2 gene-level CNV as one more cohort, PANCAN-ATLAS, under the same vector + index schema. GDC files and documents are untouched, and the existing loader and query helpers work on the new cohort without changes: download_pancan_atlas.py: fetches the four Xena matrices and records their SHA-256 in a manifest build_pancan_atlas_collections.py: builds loader-compatible collections; --verify checks the written vectors against the source matrices The view yields 9,355 samples with all three layers and an iCluster label (27 classes), all of them also present in the GDC view. It is built locally: the published database dump is unchanged. Query collection. A new queries/ folder collects AQL queries and small adapters that turn the database into the inputs of downstream analyses: a task takes its slice of the graph through one declarative query, instead of rebuilding its prior from static dumps. Four short examples show the data model at work; the extraction queries and their adapters export: typed relational subgraphs, as triples for link prediction and knowledge-graph embedding models gene-anchored multi-scale neighbourhoods (pathway, GO, disease), as node and edge lists for heterogeneous graph neural networks feature × sample omic matrices in the Xena file layout, with sample labels, for sample-level models All queries only read the database; interaction networks outside the backbone can be joined on the same identifiers. Database schema documentation. DATABASE.md describes collections, document fields and the AQL join pattern, checked against the published instance. What's new in v1.1.1 Consistent database name. Every script now defaults to the same database, PKT_main, and scripts that only read fail with an explicit message instead of silently creating an empty one; --db selects a differently named instance. Database dump and restore. scripts/db_dump.py creates, verifies, restores and downloads a dump of the materialised instance (semantic backbone and quantitative layers together), so the results can be reproduced without rebuilding the corpus. Every command counts documents and refuses to report success on a mismatch. The published dump lives at https://huggingface.co/datasets/johndef64/KG-TransomicNet: python scripts/db_dump.py download --out ./kg-dump python scripts/db_dump.py restore --input ./kg-dump --db PKT_main --create What's new in v1.1.0 Performance evaluation suite. A new scripts/benchmark/ package measures the data model as a storage and retrieval design, with all raw measurements shipped as CSV under results/benchmark/: bench_storage.py: on-disk footprint per collection bench_query.py: latency of the three retrieval modalities, with an optional secondary-index ablation bench_ingestion.py: transaction-size sweep and ingestion scaling bench_schema.py: the same measurements materialised under four alternative storage layouts (measurement-as-node, measurement-as-property, knowledge graph plus external matrix, and the proposed vector + index), compared on ingestion time, size and three access patterns Whole-corpus build driver. scripts/benchmark/run_all_cohorts.py runs download → build → load one project at a time and removes each project's raw downloads and intermediate JSON after a successful load, keeping peak disk at roughly one cohort's working set plus the growing database (~15 GB instead of ~29 GB). It records per-project timings and storage deltas, is resumable, and verifies that documents actually landed in the database rather than trusting process exit codes. Corrected corpus statistics. The quantitative layer counts are now measured on the materialised instance rather than derived from the upstream phenotype tables, which under-count TARGET projects and over-count CNV samples. The corrected totals are substantially larger than those reported in v1.0.0. Pipeline fixes. Three defects that prevented the full corpus from being ingested: load_omics_collections_to_arangodb.py: the target database is now passed explicitly by the build driver instead of silently falling back to the default build_omics_collections.py: disease_type.project is parsed tolerantly; it arrives as a serialised list for projects with several disease types but as a bare string when there is only one, which previously raised and aborted the whole project build_omics_collections.py: NumPy scalars left in documents by pandas are now serialised correctly; whether a project tripped this depended on the dtypes inferred from its own files Corrected disk requirement. The full build needs far less space than previously stated: see Requirements below. What this release contains Framework code to build, load and query the property graph in ArangoDB Curated mapping tables (data/mappings/) linking omics identifiers to knowledge-graph entities Three reproducible use cases: (UC1) predicate-stratified mRNA coherence, (UC2) multi-layer discordance classification, (UC3) phenotype-anchored traversal Benchmark suite and the raw measurements behind the reported performance figures Data artifacts The materialised database is distributed as a Hugging Face dataset due to size: https://huggingface.co/datasets/johndef64/KG-TransomicNet. It is an ArangoDB dump covering the semantic backbone and all five quantitative layers; restore it with scripts/db_dump.py restore. This release archives the code and mappings needed to build and query it. Semantic layer (ontology backbone) Source: PheKnowLator v3.0.2, instance-based OWL-NETS build (Zenodo 10689968), from the OBO Foundry and the Relation Ontology Scale: 780,753 entities, 11,082,103 typed relations Quantitative layer Five omic modalities, 16,938 distinct samples across 42 projects, measured on the materialised instance: | Layer | Platform | Projects | Samples | Features/cohort | |---|---|---:|---:|---:| | Transcriptomics | STAR TPM | 42 | 15,433 | 60,660 | | CNV (gene-level) | ASCAT3 | 33 | 10,632 | 60,623 | | miRNA | miRNA-Seq | 38 | 13,403 | 1,881 | | Proteomics | RPPA (TCPA) | 32 | 7,904 | 487 | | Methylation | Illumina HM27 | 13 | 3,137 | 27,578 | Layer availability is uneven: 11 projects carry all five layers, 22 carry four, 8 carry two or fewer. The complete instance occupies 11.4 GB: 2.4 GB semantic backbone and 9.0 GB of quantitative layers holding 1.70 × 10⁹ measurements in 50,509 per-sample vector documents plus 158 cohort index documents. Quantitative data are obtained from UCSC Xena and the GDC Data Portal (public); this release ships the mappings, not the raw omics files. Requirements Python ≥ 3.10 A running ArangoDB instance ≥ 3.11 Disk: the finished database is 11.4 GB; budget ~15 GB for a build using the per-project driver, or ~29 GB building everything in one pass License MIT

View source

Similar papers

#computer vision Conference Aug 2008

Scrum in a Multiproject Environment: An Ethnographically-Inspired Case Study on the Adoption Challenges

Agile methods continue to gain popularity. In particular, the Scrum method appears to be on the verge of becoming a de-facto standard in the industry, leading the so called Agile movement. While there are success stories and recommendations, there is little scientifically valid evidence of the challenges in the adoptio...

A. Marchenko, P. Abrahamsson · 59 citations · ⚡11
#computer vision Open access Sep 2012

Making the leap to a software platform strategy: Issues and challenges

A comprehensive taxonomy of the challenges faced when a medium-scale organization decided to adopt software platforms is provided, namely: business challenges, organizational challenges, technical challenges, and people challenges.

Yaser Ghanam, F. Maurer, P. Abrahamsson · 41 citations · ⚡3
#machine learning Open access Mar 2024

Integration of molecular coarse-grained model into geometric representation learning framework for protein-protein complex property prediction

MCGLPPI, a novel geometric representation learning framework that combines graph neural networks (GNNs) with the MARTINI molecular coarse-grained (CG) model to predict overall PPI properties accurately and efficiently, offers an effective and efficient solution for PPI overall property predictions.

Yang Yue, Shu Li, Yihua Cheng et al. · 15 citations

PepPCBench is a Comprehensive Benchmarking Framework for Protein-Peptide Complex Structure Prediction

PepPCBench enables a robust evaluation of PFNN-based methods and supports their continued development for peptide-protein structure prediction, and highlights the influence of peptide length, conformational flexibility, and training set similarity on prediction accuracy.

Si-Long Zhai, Huifeng Zhao, Ji-Ke Wang et al. · 13 citations · ⚡1
#machine learning Open access Sep 2025

Unified and explainable molecular representation learning for imperfectly annotated data from the hypergraph view

OmniMol is presented, a framework using hypergraphs to improve predictions of molecular properties, addressing challenges of imperfect data annotation and enhancing model explainability, and achieves state-of-the-art performance in properties prediction.

Bowen Wang, Junyou Li, Donghao Zhou et al. · 11 citations

Related blog posts

Microsoft Research Blog Jul 13, 2026

Verifying Rust cryptography in SymCrypt, from standards to code

Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.