Toward Imbalanced Molecular Property Regression: A Benchmark Study and Interval-Aware Mixture of Experts
Abstract
Molecular property prediction is a key task in AI-driven drug discovery, yet the prevalence and impact of label imbalances in molecular property regression remain poorly understood. Through a systematic benchmark of widely used molecular property data sets, we show that target values are often highly imbalanced and that prediction errors are consistently concentrated in sparsely represented regions of the label space. We further evaluate representative imbalance-learning approaches developed for general regression tasks and find that their effectiveness on molecular data sets is limited. To address this challenge, we propose interval-aware mixture-of-experts (IA-MoE), a plug-in framework that partitions the continuous target space into intervals and promotes expert specialization across different regions of the label distribution. IA-MoE can be seamlessly integrated with diverse molecular encoders without modifying their underlying architectures. Experiments on five molecular property benchmarks and four graph neural network backbones demonstrate that IA-MoE consistently improves the predictive performance in underrepresented regions while maintaining or improving overall accuracy. Our findings establish label imbalance as an important challenge in molecular property regression and highlight interval-aware expert learning as an effective strategy for addressing this challenge.