Genome-wide association studies (GWAS) have identified numerous variant-trait associations; yet, assigning effector genes to GWAS loci remains challenging. Similarity-based machine-learning methods, such as PoPS, prioritize effector genes from shared functional profiles among trait-relevant genes. These models assign a prioritization score for each gene and nominate a single effector gene within a GWAS locus. However, the scores provide limited insight into why a gene was prioritized or whether the nomination is biologically plausible. To address this gap, we introduce Kernelized Polygenic Priority Score, K-PoPS, a kernelized reformulation of PoPS that enables gene-centric explanations by decomposing each prediction into contributions from training genes. For each prioritized gene, K-PoPS reports top contributor genes and an anchor score that quantifies support from a user-defined set of trait-relevant genes. Across 38 Pan-UK Biobank traits, the full-feature OLS implementation underlying K-PoPS improved closest-gene enrichment relative to default PoPS for 26 of 37 evaluable traits. Across 25 traits with curated anchor sets, predictions supported by anchor scores were more enriched for closest-gene proxies than unsupported predictions. When applying to blood level apolipoprotein B, K-PoPS nominated SCARB1 over UBC gene, and further provided convincing explanations that support this prediction. Using explanation evidence, K-PoPS identified multiple plausible effector genes within a dilated cardiomyopathy locus, contrary to the parsimonious assumption. In summary, K-PoPS provides a post hoc framework for examining and interpreting GWAS effector-gene nominations.
Taotao Tan, Md. Abul Hassan Samee· bioRxiv· 0 citations
Single-cell transcriptomics has enabled systematic profiling of cellular states across ordered biological contexts, including developmental stages, treatment phases, disease progression, and anatomical compartments. A central challenge is to reconstruct trajectories that respect the directionality imposed by biology or experimental design. Existing trajectory inference methods reconstruct cell-state progressions from latent-space geometry but do not enforce external biological ordering during graph construction, yielding biologically inadmissible transitions. An emerging paradigm of optimal-transport (OT) approaches partially addresses this limitation by incorporating experimental ordering into probabilistic state-to-state correspondences, yet their pairwise formulation cannot resolve whether a given state is an intermediate state or a terminal state along a multi-step progression. In multi-timepoint settings, OT typically estimates couplings only betweenadjacent timepoints and then chains these locally solved couplings to approximate long-range trajectories without a global optimization across all conditions simultaneously. Here we present SPARC, a graph-based optimization framework that quantifies similarity in a shared high-dimensional latent space and reconstruct directional trajectories under biological constraints. Global shortest-path optimization over this graph yields progression routes, from which SPARC derives path-based pseudotime identifies bottlenecks clusters, and detects gene temporal behavior. SPARC was evaluated across three complementary settings representing distinct trajectory-inference challenges. Its application to paired primary and lung metastatic osteosarcoma samples allows us to be the first to propose a “cross-organ bone-like microenvironment” hypothesis, in which osteoclastogenic signaling establishes a bone-like remodeling niche within the pulmonary metastatic lesion that promotes osteoclast differentiation and activity. The findings are independently recoverable in human osteosarcoma Visium HD spatial transcriptomics.
Shifeng Wu, W. C. Walker, Carolyn A. Martin et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.