On the Transferability Between Extreme Multi-Label and Hierarchical Text Classification
Abstract
Extreme multi-label classification (XML) and hierarchical text classification (HTC) address closely related multi-label prediction problems, but have largely developed as separate research areas. XML focuses on very large label spaces and typically evaluates ranked label lists, while HTC assumes a human-curated label hierarchy and commonly reports classification-based F1 scores. This paper studies the transferability of representative methods across these two settings. We evaluate XML models on HTC benchmarks by flattening the human-defined hierarchy, and HTC models on XML benchmarks by inducing synthetic hierarchies. To make the comparison meaningful across both communities, we report classification metrics, ranking metrics, and R-Precision as a label-cardinality-aware bridge between them. Our results show a clear asymmetry: XML methods transfer well to HTC datasets and are often competitive with specialised HTC models, especially on ranking-based metrics. In contrast, HTC methods struggle on XML datasets, either due to scalability limits or substantially lower performance. These findings suggest that XML methods may be considered strong baselines for HTC, and that future HTC models should also be evaluated for scalability and ranking performance. Our code is available at https://github.com/FloHauss/XMC_HTC.