Multi-Modal RGB–Depth Image Segmentation Using Feature Fusion
Abstract
Image segmentation remains a challenging task, particularly in complex environments where visual information from RGB images alone is often insufficient. Factors such as poor lighting, occlusions, and background clutter can significantly degrade segmentation performance. To address these limitations, multi-modal approaches that incorporate additional data sources, such as depth information, have gained increasing attention. However, effectively combining different modalities in a meaningful way is still an open problem. In this paper, a novel multi-modal segmentation framework based on a cross-attention mechanism was presented that enables more effective interaction between RGB and depth features. Instead of relying on simple fusion strategies, the proposed method allows each modality to guide the feature selection process of the other, leading to more informative and discriminative representations. The network follows a dual-branch design, where features are first extracted independently and then fused through the proposed attention module. This paper evaluates the proposed approach on standard benchmark datasets and compares it with several baseline methods, including single-modality and early-fusion models. The results show consistent improvements in segmentation accuracy, particularly in challenging scenarios.