Jul 2026· GECCO Companion· pp. 1166-1174· 0 citations· 28 references
Computer Science
TL;DR
A clustering-based knowledge base is proposed that captures positive-class structure before search begins and uses it to focus both initialization and neighborhood exploration in Moca-I, a multi-objective local-search algorithm that mines interpretable rule sets by trading off minority-class recall, precision, and complexity.
Abstract
Mining classification rules for the minority class is challenging not only because positive examples are rare, but because they concentrate in small, geometrically irregular subregions of feature space that unguided search methods systematically miss. We propose a clustering-based knowledge base that captures positive-class structure before search begins and uses it to focus both initialization and neighborhood exploration in Moca-I, a multi-objective local-search algorithm that mines interpretable rule sets by trading off minority-class recall, precision, and complexity. Rather than sampling attribute conditions blindly from the full discretized space, the knowledge base clusters positive-class instances, extracts per-cluster attribute ranges, and aligns them with the algorithm's discretization—seeding the initial archive with minority-class-informed prototypes and restricting neighborhood operators to locally relevant regions. We instantiate this framework with two clustering methods: Self-Organizing Maps (MOCA-ISOM), which preserve the topological structure of the positive-class manifold, and K-Means (MOCA-IKM), a centroid-based baseline. Evaluated on 19 imbalanced benchmark datasets, MOCA-ISOM achieves statistically significant F-measure improvements on eight datasets and Hypervolume improvements on nine, with gains up to +17% absolute on high-dimensional data. MOCA-IKM shows comparable average performance but exhibits five statistically significant degradations.
Standard Random Forest algorithms assume crisp class boundaries and precise feature values, limitations that become critical when dealing with ambiguous or overlapping data patterns common in imbalanced datasets. This paper presents Fuzzy Random Forest (FRF), a novel ensemble method that integrates fuzzy set theory int...
James Omusula Atsali· Asian Journal of Probability...· 0 citations
Text categorization remains a challenging task due to the inherent ambiguity of natural language, class overlap, and imbalanced topic distributions. Traditional Fuzzy C-Means (FCM) clustering, although widely used for soft text classification, is highly sensitive to initialization and tends to favor dense or majority c...
Michael Loki, Agnes Mindila, W. Mwangi· Scientific Reports· 0 citations
Evaluating the performance of Greedy K-center across a variety of metric spaces shows that mapping unlabeled instances into a predictive probability space and weighting the result by entropy often dominates the other options for active learning selection with Greedy K-center.
Data-Informed Centroid Splitting (DICS), a clustering-based framework that constructs a compact and informative set of candidate splits using data-driven priors, significantly reduces the split search space for classification tasks while preserving predictive performance.
Saifur Rahman Mazumder, Feng Yu· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.