Skip to content
Open access

MatUQ: a benchmark for uncertainty-aware out-of-distribution materials property prediction with graph neural networks

Aug 2026 · npj Computational Materials · 0 citations

TL;DR

This work introduces MatUQ, a benchmark built on structure-aware Smooth Overlap of Atomic Positions Leave-One-Cluster-Out (SOAP-LOCO) splitting, together with a training protocol that combines Deep Evidential Regression (DER) with dropout regularization, for evaluating GNN reliability under structural distribution shifts.

Abstract

Reliable uncertainty quantification (UQ) for graph neural networks (GNNs) under out-of-distribution (OOD) shifts remains insufficiently characterized in materials discovery. Existing benchmarks based on random splits can overestimate model reliability by underrepresenting structural extrapolation challenges. Here we introduce MatUQ, a benchmark built on structure-aware Smooth Overlap of Atomic Positions Leave-One-Cluster-Out (SOAP-LOCO) splitting, together with a training protocol that combines Deep Evidential Regression (DER) with dropout regularization, for evaluating GNN reliability under structural distribution shifts. Through systematic experiments spanning six materials datasets, twelve GNN architectures, and eight UQ strategies, we find that predictive accuracy and uncertainty quality are distinct capabilities that tend to decouple under OOD evaluation, and that uncertainty-metric leadership is largely non-transferable across datasets and target properties. We further find that as training data become scarce, the optimal strategy shifts from evidential-containing hybrids toward pure ensemble variance across all evaluated architectures, with the architecture holding the distributional optimum shifting correspondingly. Standalone evidential regression rarely attains per-model optima and benefits from hybrid pairing only under specific combinations of data density, inductive bias, and target metric. Monte Carlo dropout is less effective as a standalone uncertainty estimator under the tested configurations but can contribute within hybrid schemes. Overall, MatUQ provides a more rigorous benchmark for assessing uncertainty-aware GNNs in OOD materials discovery.

Read PDF

Similar papers

Book Open access Aug 2026

Learning Robust Hypergraph Embeddings for Distribution-Free Uncertainty Quantification

Contrastive Conformal HGNN (CCF-HGNN) is proposed that accounts for uncertainty in hypergraph-based models by explicitly regularizing on the hypergraph structure for guaranteed and robust uncertainty estimates.

Akash Choudhuri, Bijaya Adhikari · 0 citations
Preprint Jul 2026

When does distribution shift break graph neural networks calibration?

Graph neural networks (GNNs) are increasingly deployed in real-world applications where distribution shift is un-avoidable. However, how such shifts affect model calibration, defined as the agreement between predictive confidence and actual accuracy, remains poorly understood, and existing graph calibration methods typically rely on labeled validation data from the deployment distribution. In this work, I present the first closed-form theoretical characterization of GNN calibration under distribution shift. I show that calibration is governed by a single scalar quantity that explicitly depends on structural changes between the source and target graphs, as well as feature quality. This characterization precisely identifies when a model becomes over-confident, under-confident, or remains calibrated, and directly yields the optimal temperature scaling strategy. I further extend the analysis to graph convolutional networks with symmetric normalization, multi-class classification, and covariate shift, and derive a theoretical upper bound on the expected calibration error. My analysis also reveals that, under homogeneous distribution shift, a single global temperature is theoretically optimal, providing a principled explanation for why more complex node-wise recalibration methods offer no additional benefit. Building on these theoretical insights, I propose STAC, a source-free, label-free calibration method. Experiments on synthetic benchmarks demonstrate substantial calibration improvements, while evaluations on five real-world graph datasets show that reliable calibration without target labels remains challenging despite the strong predictive power of the theory.

Abderaouf Bahi · 0 citations
Jul 2026

OpenRTAG: A Comprehensive Benchmark for Robust Text-Attributed Graph Learning under Data Quality Degradation

OpenRTAG provides a standardized testbed for understanding robustness in TAG learning under realistic low-quality settings and systematically evaluates scenario validity and model sensitivity, compares traditional GNNs, LLM-GNNs, and a representative GFM, and investigates the effectiveness, efficiency, and robustness of scenario-matched baselines.

Yu-Ze Dai, Zhi-Han Zhang, Yan Zhao et al. · 0 citations
#machine learning Preprint Aug 2026

ToxLens: A Reproducible Graph-Learning Framework for Leakage-Aware, Uncertainty-Calibrated Molecular Toxicity Prediction

Molecular toxicity prediction is increasingly used to prioritise compounds before experimental testing, but conventional benchmark performance can overstate practical utility when structurally related molecules occur across training and test folds. We introduce ToxLens, a reproducible multi-task graph-learning framework for 11 toxicity endpoints spanning Ames mutagenicity, acute oral toxicity, hERG inhibition, and Tox21 nuclear-receptor and stress-response assays. The workflow combines conservative chemical curation, sphere-exclusion filtering, a leakage-aware UMAP-HDBSCAN split, parallel graph and global-feature encoders joined by late concatenation, temperature-scaled Monte Carlo dropout with conformal-style prediction sets, applicability-domain analysis, and SHAP-guided toxicophore discovery with occlusion controls. On the leakage-controlled test fold, a five-seed soft-voting ensemble achieved a Matthews correlation coefficient score of 0.44, an area under the receiver operating characteristic curve score of 0.83, and an area under the precision-recall curve score of 0.58. It exceeded four ECFP4-based shallow baselines on all 11 endpoints under the same split and validation-based threshold-selection protocol. Controlled ablations showed that the global pathway was important, whereas late concatenation outperformed the tested gated and feature-wise linear modulation fusion variants. Conformal-style prediction sets revealed substantial endpoint-specific variation in set efficiency, and discrimination and calibration improved with similarity to the training domain. Retraining on fixed published Tox21 Challenge and TDA folds produced competitive, but not uniformly state-of-the-art, performance. SHAP-guided occlusion and consensus subgraph mining yielded model-derived structural hypotheses, 44 of which contained at least one occurrence that passed the predefined counterfactual criteria.

Magnus H. Strømme, A. D. de Sá, David B. Ascher · 0 citations
Jul 2026

HeAD-CP: Heterophily-Aware Diffused Conformal Prediction Sets for Graph Neural Networks

HeAD-CP is proposed, a family of node-wise diffusion variants whose coefficients are determined by a label-free local-homophily estimate derived from the GNN softmax, which are most effective at extreme heterophily, intermediate heterophily, and moderate-to-high homophily, respectively, and all preserve the marginal coverage guarantee.

P. Lam, Anh Thai Nguyen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.