Skip to content

Rigorous validation of graph-based network analysis reappraises biologically coherent ADHD-associated transcriptomic modules in peripheral blood

Oct 2026 · PLoS ONE · 9 references
Bioinformatics and Genomic Networks

Abstract

Background Graph attention networks (GATs) are increasingly applied to transcriptomic data because they integrate gene-network structure while producing attention weights that are often interpreted as indicators of biological importance. However, whether attention-derived explanations reliably reflect biologically meaningful signals has received little systematic evaluation, particularly in the small, heterogeneous cohorts common in psychiatric transcriptomics. Methods We systematically re-evaluated a GAT using peripheral blood RNA-sequencing data from 76 individuals (39 ADHD, 37 controls), including 16 discordant monozygotic twin pairs. To maximize analytical rigor, we implemented a leakage-aware pipeline, incorporating twin-aware group-stratified cross-validation, fold-wise gene selection, corrected transcript-to-gene mapping, and validation through repeated cross-validation, permutation testing, three complementary feature-importance methods, orthogonal differential-expression and pathway analyses, and data-quality controls. Results Across 20 repeated cross-validation splits, GAT achieved a slightly higher mean AUC than a graph convolutional network (mean AUC 0.624 vs. 0.599) and outperformed five classical machine-learning models on a prespecified split. However, GAT performance was not statistically distinguishable from a rigorously matched permutation-derived null distribution (p = 0.327), while an exploratory sample-size calculation indicated that roughly twice the current sample size would be needed to detect an effect this size at conventional power. Independent validation analyses converged on the same conclusion: attention-derived gene rankings were significantly negatively correlated with both SHAP and permutation importance, whereas differential-expression analyses---including discordant twin-pair comparisons---and pathway enrichment identified no reproducible biological signals after multiple-testing correction. Recovery of established co-expression modules and the absence of blood-cell marker differences supported pipeline validity. Conclusions Graph-based models may show modest predictive gains, but predictive performance and biological interpretation are distinct scientific questions. Our findings demonstrate that attention-derived importance should be regarded as hypothesis-generating rather than mechanistic evidence unless independently validated, illustrating why rigorous, multi-level validation is essential before biological conclusions are drawn from graph neural networks in transcriptomic research.

View source

Similar papers

#computer vision Conference Aug 2008

Scrum in a Multiproject Environment: An Ethnographically-Inspired Case Study on the Adoption Challenges

Agile methods continue to gain popularity. In particular, the Scrum method appears to be on the verge of becoming a de-facto standard in the industry, leading the so called Agile movement. While there are success stories and recommendations, there is little scientifically valid evidence of the challenges in the adoptio...

A. Marchenko, P. Abrahamsson · 59 citations · ⚡11
#computer vision Open access Sep 2012

Making the leap to a software platform strategy: Issues and challenges

A comprehensive taxonomy of the challenges faced when a medium-scale organization decided to adopt software platforms is provided, namely: business challenges, organizational challenges, technical challenges, and people challenges.

Yaser Ghanam, F. Maurer, P. Abrahamsson · 41 citations · ⚡3
#machine learning Open access Mar 2024

Integration of molecular coarse-grained model into geometric representation learning framework for protein-protein complex property prediction

MCGLPPI, a novel geometric representation learning framework that combines graph neural networks (GNNs) with the MARTINI molecular coarse-grained (CG) model to predict overall PPI properties accurately and efficiently, offers an effective and efficient solution for PPI overall property predictions.

Yang Yue, Shu Li, Yihua Cheng et al. · 15 citations

PepPCBench is a Comprehensive Benchmarking Framework for Protein-Peptide Complex Structure Prediction

PepPCBench enables a robust evaluation of PFNN-based methods and supports their continued development for peptide-protein structure prediction, and highlights the influence of peptide length, conformational flexibility, and training set similarity on prediction accuracy.

Si-Long Zhai, Huifeng Zhao, Ji-Ke Wang et al. · 13 citations · ⚡1
#machine learning Open access Sep 2025

Unified and explainable molecular representation learning for imperfectly annotated data from the hypergraph view

OmniMol is presented, a framework using hypergraphs to improve predictions of molecular properties, addressing challenges of imperfect data annotation and enhancing model explainability, and achieves state-of-the-art performance in properties prediction.

Bowen Wang, Junyou Li, Donghao Zhou et al. · 11 citations

Related blog posts

Microsoft Research Blog Jul 13, 2026

Verifying Rust cryptography in SymCrypt, from standards to code

Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.