Skip to content
Conference Open access

Improving HDBFormer with VRL for RGB-D Semantic Segmentation

2026 · ITM Web of Conferences · 0 citations · 7 references

Abstract

In RGB-D semantic segmentation, fusing the RGB branch with the depth-learning branch supports fine-grained semantic interpretation in clutteblue indoor scenes. In current computer vision practice, Transformer architectures are used almost by default, and cross-modal fusion settings follow the same trend. HDBFormer is often treated as a leading approach for RGB-D semantic segmentation. With its heterogeneous dual-branch layout, the method cuts the depth-branch computation substantially while keeping accuracy at a competitive level, yet some constraints remain. Previous analyses have demonstrated that as Transformer network depth increases, value embedding gradually loses content information during repeated attention stacking processes. During cross-modal interactions, geometric feature fading may occur earlier than expected (sometimes even before high-level semantics are fully determined). To mitigate these issues, this study incorporates a value residual learning (VRL) mechanism into the attention module. Through design, VRL enhances residual reinforcement for cross-layer value representation, which helps stabilize deep features and maintain geometric structure during RGB-depth interaction. The NYU- Depth-V2 benchmark results demonstrate that mIoU metric shows significant improvement with increasing VRL, while the number of parameters and computational load remain largely unchanged. Based on these research findings, this study provides practical guidance for optimizing efficient RGB-D semantic segmentation models without incurring additional costs.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.