Skip to content

Author

Tae‐Sik Park

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#graph neural networks Open access Sep 2026

ToxBench: a leakage-audited multi-task benchmark for predictive toxicology with calibration, uncertainty, and applicability domain analysis

Machine learning models for toxicity prediction are routinely evaluated using random train/test splits, allowing structurally similar compounds to appear in both sets and inflating reported performance metrics. A rigorous benchmark quantifying this overestimation alongside calibration, uncertainty, and applicability domain analyses is needed to guide practitioners in selecting and trusting toxicity prediction models. We present ToxBench, a leakage-audited benchmark comprising three toxicology datasets (Tox21, 7,538 compounds, 12 tasks; ClinTox, 1,379 compounds, 2 tasks; SIDER, 1,350 compounds, 27 tasks) processed through a transparent standardization pipeline with explicit reporting of conflicting-label removals. Four model classes were evaluated: Random Forest, XGBoost, MLP, and Graph Neural Network (GNN), each trained under random and Bemis-Murcko scaffold-based splits across five independent seeds (120 experimental conditions). Analyses included post-hoc probability calibration, ensemble-based uncertainty quantification, nearest-neighbor applicability domain analysis, and scaffold-level error analysis. Scaffold splitting consistently reduced AUROC by 0.057–0.079 points across all model classes on Tox21 (mean drop: 0.070) and by 0.031–0.035 points on SIDER for three of four models, demonstrating systematic performance overestimation under random splitting. ClinTox showed reversed performance ordering due to small dataset size and extreme class imbalance, with high seed-to-seed variance (± 0.085–0.160) confirming results are dominated by sampling noise. Post-hoc calibration reduced ECE by 67–68% on ClinTox, confirming reliable probability estimates are achievable with simple correction. On Tox21 and SIDER, raw Random Forest predictions were already well-calibrated (ECE 0.018 and 0.057 respectively), with post-hoc methods providing no additional benefit. Scaffold splitting consistently worsened calibration across all three datasets. Applicability domain analysis revealed a consistent positive relationship between nearest-neighbor Tanimoto similarity and predictive reliability, with low-similarity compounds showing notably reduced AUROC under scaffold splitting. Ensemble uncertainty estimates showed promising utility, with confidence-based filtering improving AUROC by up to 0.025 points on SIDER. Random splitting overestimates performance by 0.057–0.079 AUROC points on Tox21 and 0.031–0.035 on SIDER. ToxBench provides a transparent, reproducible, leakage-audited benchmark supporting scaffold splitting, multi-seed evaluation, and applicability domain filtering as standard practice in toxicity prediction benchmarking. Rather than assuming scaffold splitting is leakage-free, we audit residual structural similarity directly: scaffold splitting substantially reduces, but does not eliminate, train–test analog leakage (e.g., the fraction of Tox21 test compounds with a nearest-neighbour Tanimoto > 0.6 to training falls from 43.9% under random splitting to 12.1% under scaffold splitting), and we quantify this residual and provide an applicability-domain filter to flag it.

Kartic, Yeeun Seo, Sanggyun Yi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.