KG-TransomicNet: Semantic–Quantitative Property Graph of PheKnowLator and Multi-Omics
Abstract
KG-TransomicNet Version 1.1.2 · Software · Open · MIT De Filippis, Giovanni Maria · Rinaldi, Antonio Maria A semantic–quantitative property-graph framework that couples the PheKnowLator biomedical knowledge graph with per-sample multi-omics measurements in ArangoDB, enabling trans-omic analysis and knowledge-based reasoning over TCGA/TARGET cohorts via AQL. What's new in v1.1.2 PanCanAtlas view. A new scripts/pancan_atlas/ package adds the harmonised TCGA PanCanAtlas data (EB++ RNA-seq, batch-adjusted miRNA, iCluster subtype labels) and the pan-cancer GISTIC2 gene-level CNV as one more cohort, PANCAN-ATLAS, under the same vector + index schema. GDC files and documents are untouched, and the existing loader and query helpers work on the new cohort without changes: download_pancan_atlas.py: fetches the four Xena matrices and records their SHA-256 in a manifest build_pancan_atlas_collections.py: builds loader-compatible collections; --verify checks the written vectors against the source matrices The view yields 9,355 samples with all three layers and an iCluster label (27 classes), all of them also present in the GDC view. It is built locally: the published database dump is unchanged. Query collection. A new queries/ folder collects AQL queries and small adapters that turn the database into the inputs of downstream analyses: a task takes its slice of the graph through one declarative query, instead of rebuilding its prior from static dumps. Four short examples show the data model at work; the extraction queries and their adapters export: typed relational subgraphs, as triples for link prediction and knowledge-graph embedding models gene-anchored multi-scale neighbourhoods (pathway, GO, disease), as node and edge lists for heterogeneous graph neural networks feature × sample omic matrices in the Xena file layout, with sample labels, for sample-level models All queries only read the database; interaction networks outside the backbone can be joined on the same identifiers. Database schema documentation. DATABASE.md describes collections, document fields and the AQL join pattern, checked against the published instance. What's new in v1.1.1 Consistent database name. Every script now defaults to the same database, PKT_main, and scripts that only read fail with an explicit message instead of silently creating an empty one; --db selects a differently named instance. Database dump and restore. scripts/db_dump.py creates, verifies, restores and downloads a dump of the materialised instance (semantic backbone and quantitative layers together), so the results can be reproduced without rebuilding the corpus. Every command counts documents and refuses to report success on a mismatch. The published dump lives at https://huggingface.co/datasets/johndef64/KG-TransomicNet: python scripts/db_dump.py download --out ./kg-dump python scripts/db_dump.py restore --input ./kg-dump --db PKT_main --create What's new in v1.1.0 Performance evaluation suite. A new scripts/benchmark/ package measures the data model as a storage and retrieval design, with all raw measurements shipped as CSV under results/benchmark/: bench_storage.py: on-disk footprint per collection bench_query.py: latency of the three retrieval modalities, with an optional secondary-index ablation bench_ingestion.py: transaction-size sweep and ingestion scaling bench_schema.py: the same measurements materialised under four alternative storage layouts (measurement-as-node, measurement-as-property, knowledge graph plus external matrix, and the proposed vector + index), compared on ingestion time, size and three access patterns Whole-corpus build driver. scripts/benchmark/run_all_cohorts.py runs download → build → load one project at a time and removes each project's raw downloads and intermediate JSON after a successful load, keeping peak disk at roughly one cohort's working set plus the growing database (~15 GB instead of ~29 GB). It records per-project timings and storage deltas, is resumable, and verifies that documents actually landed in the database rather than trusting process exit codes. Corrected corpus statistics. The quantitative layer counts are now measured on the materialised instance rather than derived from the upstream phenotype tables, which under-count TARGET projects and over-count CNV samples. The corrected totals are substantially larger than those reported in v1.0.0. Pipeline fixes. Three defects that prevented the full corpus from being ingested: load_omics_collections_to_arangodb.py: the target database is now passed explicitly by the build driver instead of silently falling back to the default build_omics_collections.py: disease_type.project is parsed tolerantly; it arrives as a serialised list for projects with several disease types but as a bare string when there is only one, which previously raised and aborted the whole project build_omics_collections.py: NumPy scalars left in documents by pandas are now serialised correctly; whether a project tripped this depended on the dtypes inferred from its own files Corrected disk requirement. The full build needs far less space than previously stated: see Requirements below. What this release contains Framework code to build, load and query the property graph in ArangoDB Curated mapping tables (data/mappings/) linking omics identifiers to knowledge-graph entities Three reproducible use cases: (UC1) predicate-stratified mRNA coherence, (UC2) multi-layer discordance classification, (UC3) phenotype-anchored traversal Benchmark suite and the raw measurements behind the reported performance figures Data artifacts The materialised database is distributed as a Hugging Face dataset due to size: https://huggingface.co/datasets/johndef64/KG-TransomicNet. It is an ArangoDB dump covering the semantic backbone and all five quantitative layers; restore it with scripts/db_dump.py restore. This release archives the code and mappings needed to build and query it. Semantic layer (ontology backbone) Source: PheKnowLator v3.0.2, instance-based OWL-NETS build (Zenodo 10689968), from the OBO Foundry and the Relation Ontology Scale: 780,753 entities, 11,082,103 typed relations Quantitative layer Five omic modalities, 16,938 distinct samples across 42 projects, measured on the materialised instance: | Layer | Platform | Projects | Samples | Features/cohort | |---|---|---:|---:|---:| | Transcriptomics | STAR TPM | 42 | 15,433 | 60,660 | | CNV (gene-level) | ASCAT3 | 33 | 10,632 | 60,623 | | miRNA | miRNA-Seq | 38 | 13,403 | 1,881 | | Proteomics | RPPA (TCPA) | 32 | 7,904 | 487 | | Methylation | Illumina HM27 | 13 | 3,137 | 27,578 | Layer availability is uneven: 11 projects carry all five layers, 22 carry four, 8 carry two or fewer. The complete instance occupies 11.4 GB: 2.4 GB semantic backbone and 9.0 GB of quantitative layers holding 1.70 × 10⁹ measurements in 50,509 per-sample vector documents plus 158 cohort index documents. Quantitative data are obtained from UCSC Xena and the GDC Data Portal (public); this release ships the mappings, not the raw omics files. Requirements Python ≥ 3.10 A running ArangoDB instance ≥ 3.11 Disk: the finished database is 11.4 GB; budget ~15 GB for a build using the per-project driver, or ~29 GB building everything in one pass License MIT