Network Organization Of Genetic Susceptibility to Disease
Abstract
The file dohai_thesis_core_v1.zip contains the analysis scripts and associated input files used for the doctoral thesis titled “Network Organization of Genetic Susceptibility to Disease”. The scripts are organized according to the analytical workflow described in the Methods chapter. The workflow includes harmonization of gene trait associations from Open Targets Genetics and OMIM, profiling of human protein coding genes, analysis of gene pleiotropy and phenotype polygenicity, construction and characterization of biochemical networks, assessment of causal gene coverage, degree distribution analysis, K shell decomposition, analysis of causal gene homophily and network proximity, identification of shortest path nodes, functional enrichment analysis, and evaluation of functional gene sets as positive controls. The functional positive controls include Gene Ontology biological process, molecular function and cellular component annotations, KEGG pathways, and CORUM protein complexes. The corresponding analyses were performed using the HuRI, BioPlex 3.0 and Recon3D networks. Scripts are grouped into numbered folders following the order of the Methods chapter. Each folder contains a README describing the purpose of the scripts, required inputs, expected outputs and recommended execution order. Some analyses depend on outputs generated by preceding stages of the workflow. The archive excludes utilities used solely to edit or generate thesis documents. It contains only the scientific analysis scripts and their supporting input files. The supplementary datasets accompanying this thesis are provided in dohai_thesis_supplementary_data.zip. The archive contains nine Excel workbooks describing the organisation of disease- and non-disease-associated causal genes across biochemical networks. These include a harmonised Open Targets–OMIM gene–phenotype atlas, network coverage, direct-interaction and proximity analyses, clustering results, functional enrichment, and functional gene sets used as positive controls to evaluate the pipeline. Supplementary Data 1 — Causal genes: Study-level associations, the harmonised gene–phenotype atlas, gene annotations and pleiotropy measures, including an analysis excluding OMIM associations. Supplementary Data 2 — Network coverage: Representation of causal genes across biochemical networks, connected-component membership and phenotype overlap between networks. Supplementary Data 3 — Direct interactions among causal genes: Direct interactions within phenotype-associated and disease-group gene sets, with significance assessed against degree-preserving randomised networks. Supplementary Data 4 — Single-neighbourhood modularity: Proximity of phenotype-associated causal genes without clustering, assessed using mean pairwise shortest-path lengths and degree-preserving randomised networks. Supplementary Data 5 — Multiple-module analysis: Gene clusters, cluster proximity tests, overlap of modular phenotypes across networks and proportions of represented causal genes belonging to significant clusters. Supplementary Data 6 — Modularity dependencies: Relationships between modularity, study count, represented causal-gene count and gene pleiotropy. Supplementary Data 7 — Functional enrichment of trait modules: Module construction using causal-gene clusters and shortest-path intermediary nodes, together with functional enrichment results. Broader enrichment outputs are retained; results selected for interpretation focus on significant Gene Ontology biological process terms with more than one overlapping gene. Supplementary Data 8 — Functional positive controls without clustering: Whole-set proximity analyses of KEGG pathways, Gene Ontology functional sets and CORUM protein complexes, compared with disease and non-disease causal-gene sets. Supplementary Data 9 — Functional positive controls with clustering: Clustering-based analyses of functional gene sets, including significant clusters, modular-gene proportions, distribution summaries and statistical comparisons with disease and non-disease causal-gene sets. Each workbook begins with a guide describing its worksheets, columns, methods and relevant thresholds. Worksheet names combine the supplementary data number with a letter, with A identifying the opening guide.