Jul 2026· International Journal of Molecular Sciences· Vol 27· 0 citations· 36 references
Medicine
TL;DR
This work establishes a standardized framework for evaluating non-coding SNP representations and offers guidance for selecting and optimizing prediction pipelines in regulatory genomics.
Abstract
Non-coding single nucleotide polymorphisms (SNPs) are key modulators of gene regulation and have been implicated in diverse complex traits and diseases. With the growing demand for accurate functional interpretation of non-coding variants, the choice of encoding strategies becomes critical in downstream predictive modeling. Despite recent advances, a systematic evaluation of encoding approaches tailored for non-coding SNPs remains lacking. To address this gap, we present a comprehensive benchmark that evaluates six representative encoding strategies, including categorical, semantic, and functional embeddings, across three quantitative trait loci (QTL)-related prediction tasks. The study encompasses nine machine learning and deep learning models and incorporates experimental controls and repeated trials to ensure robustness and reproducibility. We assess each strategy along multiple dimensions, such as interpretability, representation abundance, and computational efficiency. Rather than ranking individual methods, our analysis emphasizes the interaction between encoding strategies, model types, and preprocessing protocols, and highlights their collective influence on predictive performance. This work establishes a standardized framework for evaluating non-coding SNP representations and offers guidance for selecting and optimizing prediction pipelines in regulatory genomics.
An encyclopedia of enhancer–gene regulatory interactions in the human genome is built, revealing global properties of enhancer networks, identifying differences in regulatory complexity across genes, and improving analyses linking noncoding variants to target genes and cell types for common, complex diseases.
A. Gschwind, Kristy S. Mualim, Alireza Karbalayghareh et al.· Nature· 6 citations
This work developed a scoring-approach for AI-agents to autonomously assess AlphaGenome prediction confidence and accurately differentiate between AlphaGenome’s robust sequence-level recognition across species and its current limitations when interpreting un-fine-mapped regulatory variants.
Priya Ramarao-Milne, Suyu Ma, L. Sng et al.· bioRxiv· 0 citations
This review compares SNP detection programs such as GATK, BCFtools, FreeBayes, SAMtools, SAMtools, and DeepVariant and their algorithmic structures, namely pileup- based, haplotype-based, and machine-learning approaches and suggests that no single tool is the best.
Shikhi Baruri, Sunita Khanal· Nepal Journal of Biotechnolo...· 0 citations
Deciphering how non-coding variants perturb gene regulation is central to translating GWAS loci into mechanism, yet existing prioritization methods rarely deliver cell-type-resolved molecular effects, causal variant-to-gene attribution, or principled reasoning about combinatorial interactions. We introduce MUGO (Multi-head Genomic Optimization), an in silico perturbation framework that casts variant discovery as differentiable combinatorial optimization over genomic sequence. MUGO relaxes discrete edits into a continuous probabilistic nucleotide mask and performs gradient-based optimization in input space to identify single- or multi-variant perturbations that maximize a user-specified molecular objective under a sequence-to-signal foundation model. This formulation makes genome-scale search computationally tractable while retaining direct, cell-type-specific molecular readouts and enabling precise quantification of non-additive interaction effects. Across five modalities and seven tissues using two foundation-model backbones, MUGO consistently outperforms three baselines in both optimization efficiency and effect modulation, while preserving robustness and cell-type specificity. Finally, MUGO-prioritized variants are enriched for GWAS signals across diverse tissues and a broad spectrum of complex traits, turning foundation-model predictions into scalable, cell-type-resolved hypotheses for causal variant discovery and combinatorial regulatory mechanisms. Code and documentation are available at https://github.com/aicb-ZhangLabs/MUGO.
Si-Ying Sun, Junhao Liu, Pengcheng Xu et al.· Proceedings of the 32nd ACM...· 0 citations
This review compares convolutional, Transformer-based and graph architectures used to represent local sequence features, chromatin state and three-dimensional genome organisation to their applications to transcription-factor binding, chromatin accessibility, gene expression, non-coding variant prioritisation and regulatory-sequence design.