A multimodal BEV 3D object detection method with depth uncertainty and geometric saliency
Abstract
High-precision perception is fundamental to safe autonomous driving, and BEV-based 3D object detection via lidar-camera fusion plays a crucial role in improving detection accuracy and robustness. To address insufficient feature representation, spatial misalignment, and the limitations of static fusion strategies, this work develops a multi-modal collaborative optimization framework: (1)A depth uncertainty-guided feature modulation method is proposed, in which depth entropy and variance are jointly modeled to generate a BEV alignment confidence map, enabling adaptive enhancement and suppression of image features and effectively mitigating cross-modal alignment errors. (2) We propose geometric saliency pillar feature encoding, which enhances point cloud structural representation via point-wise saliency weighting and multi-statistic aggregation. (3)We design an adaptive interaction fusion strategy that explicitly models crossmodal consistency and discrepancy relationships, generating adaptive weights to achieve dynamic fusion of multimodal features. Experiments on the NuScenes benchmark and a self-collected T23 dataset show that mAP improves from 62.1% to 63.1% and from 68.8% to 70.1%, respectively, with NDS reaching 69.6 and 75.6. Meanwhile, mATE, mASE, and mAOE consistently decrease, demonstrating that the proposed modules provide differentiated contributions and exhibit strong synergistic effects in performance optimization.