Skip to content

Robust Multimodal Gated Fusion for 3-D Object Detection via Alignment and Denoising

Oct 2026 · IEEE Sensors Journal · Vol 26, pp. 29409-29420 · 0 citations · 49 references

Abstract

Multimodal perception integrating light detection and ranging (LiDAR) and cameras has become a key paradigm for 3-D object detection, as it leverages both geometric structure and semantic information. However, in real-world autonomous driving scenarios, calibration errors, adverse weather, and sensor degradation can introduce cross-modal spatial misalignment and unreliable features, degrading the robustness and stability of fusion-based models. To address these issues, this article proposes a robust multimodal gated fusion (RMGF) framework based on alignment and denoising, termed RMGF. The framework adopts a decoupled design with three key components. First, a graph-based spatial alignment (GSA) module geometrically calibrates camera features in the bird’s-eye-view (BEV) space to alleviate cross-modal inconsistencies. Second, a cross-modal consistency gate (CMCG) suppresses unreliable features by learning spatially adaptive weights. Third, a denoising feature refinement (DFR) module, inspired by diffusion processes, refines fused features in a residual manner to mitigate distribution shifts caused by degraded inputs. Experimental results on the nuScenes dataset show that, under identical settings, the proposed method achieves 69.4% mean average precision (mAP) and 72.1% nuScenes detection score (NDS), yielding a modest improvement while maintaining comparable performance. Further evaluations under degraded conditions demonstrate more stable performance. Overall, the proposed decoupled alignment and fusion strategy improves the stability of multimodal 3-D object detection under nonideal sensing conditions.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.