Improving HDBFormer with VRL for RGB-D Semantic Segmentation
In RGB-D semantic segmentation, fusing the RGB branch with the depth-learning branch supports fine-grained semantic interpretation in clutteblue indoor scenes. In current computer vision practice, Transformer architectures are used almost by default, and cross-modal fusion settings follow the same trend. HDBFormer is o...