Identifying core genes and exploring the omnigenic architecture of complex traits using interpretable graph neural networks
Abstract
Understanding the relationship between phenotype and genotype is a central challenge of 21st century biology. Although rare, strong mutations allow for a direct understanding of this relationship, modern genetics is puzzled by the finding that complex traits are influenced by hundreds or thousands of variants, often without a direct mechanistic link to the trait. Recently, the omnigenic model has been proposed, formulating a theoretical concept which embeds genes in a network context where few genes, the core genes, have a sizable and direct influence on a trait and the remainder of the genome, the peripheral genes, have small but non-zero indirect effects on the trait by influencing the trait’s core genes. Although appealing for its intuitive simplicity, attempts to move this theoretical concept to application have been rare and only for specific traits without generalizable conclusion. Therefore, the omnigenic model has not yet fulfilled its promises, such as the identification of prime candidates for drug development and diagnosis, and a deepened understanding of complex traits. In the first part of this thesis, a positive-unlabeled graph machine learning ensemble framework, Speos, is developed to identify core genes for a diverse set of trait groups. Speos utilizes a multimodal data suite ranging from genome-wide association studies (GWAS) and tissue-specific gene expression data to protein-protein interaction (PPI) and gene regulatory networks (GRNs) to assign each gene a consensus score, reflecting the certainty of the model ensemble regarding the genes core gene character. During training, Mendelian disorder genes serve as ground truth positive examples. The novel core genes identified by Speos are enriched for trait-matched mouse knockout genes on par with gold standard Mendelian disorder genes. Moreover, the identified core genes are differentially expressed in the presence of the respective disease and are intolerant to functional mutations, both characteristics predicted for core genes by the omnigenic model. While Mendelian disorder genes and the newly identified core genes are both heavily enriched for drug targets, only the latter is also enriched for druggable genes among the non-drug-targets, indicating the translational potential of these gene sets. Therefore, Speos can be regarded as a valuable tool for core gene discovery that is able to generalize across a wide variety of trait groups. After introducing and validating Speos, the second part of the thesis applies it to ulcerative colitis (UC), an inflammatory bowel disease (IBD), to obtain a better understanding of this traits omnigenic architecture. Core genes identified for UC are highly expressed in disease-relevant tissues and have direct, mechanistically plausible connections to the disease. Peripheral genes with a similarly high GWAS signal, on the other hand, are expressed in many tissues without connection the UC and have broader, regulatory roles. Similarly, network connections with a high importance for the prediction of core genes are predominantly PPIs and regulatory relationships in disease-relevant tissues. These patterns are strikingly different in two contrasting traits, coronary artery disease (CAD) and schizophrenia (SCZ) , indicating their trait-specificity. Furthermore, core genes show concerted up- or downregulation in response to perturbation of a large part of the genome, a behavior that is not observed for GWAS genes or randomly sampled genes, indicating that core genes are part of tightly co-regulated cellular programmes. These programmes are specific both to the trait and to cell type, indicating that the discovery of disease-causing mechanisms requires a careful selection of model cell lines for each trait. Finally, co-perturbation simulations of more than 100,000 gene pairs indicate that pairs of core genes, if perturbed jointly, lead to non-linear outcomes such as suppression and neomorphism effects more often than expected. This adds a novel layer of complexity not previously considered in the omnigenic model. The results presented in this work provide several important insights. First, it presents a computational framework for core gene identification that is general enough to be applied to a wide range of traits while simultaneously identifying bona fide core genes. This improves upon the iterative nature of previous efforts, making core gene identification more accessible and the identified core genes comparable across traits. Second, identified core genes exhibit distinct expression patterns in disease-relevant tissues that separate them from peripheral genes with a similarly strong GWAS signal. This allows subsequent research to focus their efforts on potentially promising tissues, enabling a higher resolution of disease-relevant processes. Third, this work shows that core gene sets frequently react jointly to perturbations across the genome, but the perturbagens leading to such a discriminative perturbation differ across cell lines. Thus, to find genes that exert a strong influence on core genes in vivo, and therefore have the potential to affect disease processes, cell lines or primary cell cultures should be closely matched to disease-relevant tissues and cell types. This has the potential to expand the applications of the omnigenic model beyond the compilation of core gene lists into an exploration of the genome-wide network effect. Furthermore, co-perturbation simulations indicate that joint perturbations of core genes may lead to strong non-linear effects more often than expected. Given how frequent core genes react jointly to perturbations, this indicates that the interdependencies between core genes and their joint effect on the trait must be considered, adding a novel layer to the omnigenic model. Finally, this work indicates that the previous characterizations of core genes might be insufficient and that it is more likely that a precise definition of core genes might depend on the trait, encouraging an iterative refinement of putative core gene sets by domain experts.