Skip to content
Book Open access

Clustering-Guided Knowledge Base for Multi-Objective Rule Mining on Imbalanced Datasets

Jul 2026 · GECCO Companion · pp. 1166-1174 · 0 citations · 28 references
Computer Science

TL;DR

A clustering-based knowledge base is proposed that captures positive-class structure before search begins and uses it to focus both initialization and neighborhood exploration in Moca-I, a multi-objective local-search algorithm that mines interpretable rule sets by trading off minority-class recall, precision, and complexity.

Abstract

Mining classification rules for the minority class is challenging not only because positive examples are rare, but because they concentrate in small, geometrically irregular subregions of feature space that unguided search methods systematically miss. We propose a clustering-based knowledge base that captures positive-class structure before search begins and uses it to focus both initialization and neighborhood exploration in Moca-I, a multi-objective local-search algorithm that mines interpretable rule sets by trading off minority-class recall, precision, and complexity. Rather than sampling attribute conditions blindly from the full discretized space, the knowledge base clusters positive-class instances, extracts per-cluster attribute ranges, and aligns them with the algorithm's discretization—seeding the initial archive with minority-class-informed prototypes and restricting neighborhood operators to locally relevant regions. We instantiate this framework with two clustering methods: Self-Organizing Maps (MOCA-ISOM), which preserve the topological structure of the positive-class manifold, and K-Means (MOCA-IKM), a centroid-based baseline. Evaluated on 19 imbalanced benchmark datasets, MOCA-ISOM achieves statistically significant F-measure improvements on eight datasets and Hypervolume improvements on nine, with gains up to +17% absolute on high-dimensional data. MOCA-IKM shows comparable average performance but exhibits five statistically significant degradations.

Read PDF

Similar papers

Open access Aug 2026

Fuzzy Random Forest: Integrating Fuzzy Set Theory for Enhanced Imbalanced Classification

Standard Random Forest algorithms assume crisp class boundaries and precise feature values, limitations that become critical when dealing with ambiguous or overlapping data patterns common in imbalanced datasets. This paper presents Fuzzy Random Forest (FRF), a novel ensemble method that integrates fuzzy set theory int...

James Omusula Atsali · 0 citations
Open access Jul 2026

Semantic neighborhood-aware fuzzy clustering for balanced text categorization.

Text categorization remains a challenging task due to the inherent ambiguity of natural language, class overlap, and imbalanced topic distributions. Traditional Fuzzy C-Means (FCM) clustering, although widely used for soft text classification, is highly sensitive to initialization and tends to favor dense or majority c...

Michael Loki, Agnes Mindila, W. Mwangi · 0 citations
Preprint Aug 2026

Diversity-Based Active Learning: An Evaluation of Metric Spaces for Active Learning Selection

Evaluating the performance of Greedy K-center across a variety of metric spaces shows that mapping unlabeled instances into a predictive probability space and weighting the result by entropy often dominates the other options for active learning selection with Greedy K-center.

Siddharth Chilamkur, D. Hochbaum · 0 citations
Preprint Aug 2026

DICS: Data-Informed Centroid Splitting for Decision Tree Classifiers

Data-Informed Centroid Splitting (DICS), a clustering-based framework that constructs a compact and informative set of candidate splits using data-driven priors, significantly reduces the split search space for classification tasks while preserving predictive performance.

Saifur Rahman Mazumder, Feng Yu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.