Improving Diagnostic Sensitivity in Imbalanced Oral Cancer Image Classification: A Comparative Study of CNN and Transformer Architectures
Abstract
Simple Summary Early detection of oral cancer is essential for improving patient survival, yet artificial intelligence models often struggle because available clinical image datasets are small and highly imbalanced, with relatively few malignant cases. This study investigates whether class imbalance can be mitigated through random under-sampling, random over-sampling, and medically informed image augmentation when training three representative deep learning architectures: EfficientNet, Vision Transformer, and Swin Transformer. Using a publicly available dataset of oral cavity photographs and a unified five-fold cross-validation framework, we compared these strategies using both multiclass classification and a clinically oriented malignant-risk evaluation. The results indicate that data augmentation generally improves minority-class detection while maintaining competitive overall classification performance. Nevertheless, external multicenter validation and prospective clinical studies are required before routine clinical application.