VCE-DINO: Robust Video Capsule Endoscopy Anatomy Recognition via DINO Pretraining and Imbalance-Aware Learning
Abstract
Video Capsule Endoscopy (VCE) enables noninvasive visualization of the gastrointestinal (GI) tract by capturing large-scale image streams as the capsule traverses multiple anatomical regions. Accurate recognition of major anatomical locations (e.g., mouth, esophagus, stomach, small intestine, and colon) is essential for reliable downstream diagnostic analysis. In this work, we propose VCE-DINO, a robust framework for capsule endoscopy anatomy recognition that integrates self-supervised pretraining with imbalance-aware learning. Specifically, we employ a DINOv3 Vision Transformer with Low-Rank Adaptation (LoRA) to achieve efficient adaptation while preserving strong generalization. To address class imbalance, we introduce an Imbalance-Aware Weighted Cross-Entropy Loss that stabilizes training for minority classes. In addition, we apply targeted photometric augmentations to improve robustness under the harsh lighting conditions. We evaluate our method on the large-scale Galar dataset (3.5M images), achieving a macro F1score of 0.8822 and a macro accuracy of 0.9716, outperforming a ResNet-50 baseline by 17.22% and 16.16%, respectively. These results demonstrate the effectiveness of our approach as a reliable anatomical filtering module for scalable GI diagnostic pipelines.