A Subspace Ensemble Framework for High-Dimensional Active Learning
Abstract
Supervised learning in fields such as genomics and medical imaging is often hindered by the high cost of expert data annotation. Active learning addresses this bottleneck by iteratively selecting the most informative unlabeled samples for labeling. However, in high-dimensional environments, traditional diversity-based query strategies lose their effectiveness due to the degradation of global distance metrics. To address these challenges, this paper proposes a novel framework, Active Learning via Subspace Ensembles and Similarity (ALSES). Instead of relying on global distances, ALSES constructs a similarity matrix by sampling an ensemble of random feature subspaces. The subspaces are filtered based on their discriminative power, and pairwise sample similarities are aggregated using cluster co-occurrence. This structural representation is integrated into a hybrid batch selection strategy that balances model uncertainty and data representativeness. Extensive evaluations on simulated datasets and real-world high-dimensional cancer cohorts demonstrate that ALSES consistently outperforms standard active learning baselines. The framework effectively isolates informative variables and achieves superior classification accuracy with significantly fewer labeled instances, demonstrating its robustness in complex, noisy applications.