This work developed a scoring-approach for AI-agents to autonomously assess AlphaGenome prediction confidence and accurately differentiate between AlphaGenome’s robust sequence-level recognition across species and its current limitations when interpreting un-fine-mapped regulatory variants.
Abstract
Deep learning models that predict molecular phenotypes directly from DNA sequence offer a powerful framework for interpreting genomic variation. Recently, AlphaGenome was introduced as a deep sequence-to-function architecture capable of predicting observations that historically required experiments. While the model has shown high accuracy, it was primarily evaluated on human variants scored against a reference genome. Here, we test performance on mouse data, the other species AlphaGenome was trained on although with fivefold fewer features than human (1,128 versus 5,930). We demonstrate that AlphaGenome’s predictive performance varies considerably depending on the functional task. Specifically, predicted quantitative expression effects are directionally weak and compressed roughly 100-fold relative to empirical benchmarks across both reconstructed-haplotype and single-variant regimes. In contrast, canonical splice-site disruptions are recognized with near-identical accuracy in mouse and human (AUC 0.96 versus 0.98), displaying no cross-species divergence in predicted effect magnitude. We developed a scoring-approach for AI-agents to autonomously assess AlphaGenome prediction confidence and accurately differentiate between AlphaGenome’s robust sequence-level recognition across species and its current limitations when interpreting un-fine-mapped regulatory variants. This demonstrates how GenAI innovations that are still under development can safely be harnessed by wrapping a responsible AI layer around the call to intercept flawed results, thereby adhering to international standards, such as the Australian Voluntary AI Safety Standard (VAISS).
Deciphering the regulatory consequences of sequence divergence across human evolution is essential to understanding the molecular basis of human-specific traits and disease. Although millions of derived alleles distinguish humans from great apes, only a small fraction are likely to influence human-specific traits. Previous studies have focused on regions of elevated sequence divergence, assuming that rapid evolution reflects functional adaptation, yet individual high-impact regulatory mutations evade such scans. Here, we apply sequence-to-function deep learning to predict chromatin accessibility across modern human, archaic hominin, and great ape personalized genomes, identifying lineage-specific cis-regulatory elements (linCREs) across diverse cellular contexts. Compared to conserved elements, linCREs are shorter, less pleiotropic, less conserved, and enriched in neurodevelopmental pathways. Many linCREs occur in regions with limited sequence divergence that acceleration-based approaches would overlook. We validate lineage-specific enhancer activity through luciferase reporter assays and demonstrate that a single motif-generating derived allele nominated by model interpretability tools drives a hominin-specific neurodevelopmental enhancer.
Riley J. Mangan, Nikitha Thoduguli, Dimitar Ivanov et al.· bioRxiv· 0 citations
It is demonstrated that modern EC predictors largely fail to distinguish catalytically incompetent variants from functional enzymes, and it is proposed that integrating structure-aware negative examples into both training and benchmarking is critical for developing functionally robust models in computational enzymology.
João Sartori, Ana Carolina Ramos Guimarães, Lucas de Almeida Machado· bioRxiv· 0 citations
This work establishes a standardized framework for evaluating non-coding SNP representations and offers guidance for selecting and optimizing prediction pipelines in regulatory genomics.
Hui Jin, Yihang Bao, Wenhao Li et al.· International Journal of Mol...· 0 citations
Computational predictors of RNA splicing are increasingly used to interpret genetic variants and to design synthetic genes, yet they are almost always benchmarked on endogenous human sequences closely related to their training data. Whether their performance reflects genuine recognition of splicing signals, or instead exploits statistical features of natural genomes such as conservation and exon–intron composition, remains unclear. Here we benchmark eleven splicing predictors on thousands of synthetic GFP variants that are heavily recoded and dissimilar from any training data, using long-read sequencing to measure splicing directly at each position. Despite this distribution shift, modern deep-learning predictors retained strong performance, and the resulting ranking was largely stable across position-level and construct-level benchmarks. SpliceTransformer ranked highest, followed by AlphaGenome and SpliceAI. Tools that ignore long-range sequence context performed substantially worse, largely because they assign high scores to many non-spliced positions. This ranking broadly agrees with benchmarks on endogenous variants, indicating that the leading models capture transferable, sequence-intrinsic determinants of splicing. We further provide a unified calibration that maps each predictor’s scores onto the measured fraction of spliced reads, allowing scores to be interpreted as splicing outcomes and compared directly between tools. Our results show that current deep-learning models generalise beyond natural genomes and provide a practical framework for splicing-aware sequence design.
Fernando Bellido Molías, G. Kudla· bioRxiv· 0 citations
GB.GeneUnet, an 837M-parameter transformer-based U-Net pretrained on 6 trillion tokens from multi-species genomes in OpenGenome2 is introduced, extending genomic context to 1 Mb with up to 100× inference speedup over GeneMoE, a preliminary MoE transformer baseline of similar model size pretrained on the same data.
Ning Sun, William de Vazelhes, Pan Li et al.· bioRxiv· 0 citations