Two lightweight and complementary modules to enhance voxel feature quality for NeRF-based 3D detection with consistent improvements over the NeRF-RPN baseline in both recall and precision are introduced.
Abstract
3D object detection from multi-view images has gained increasing attention as a cost-effective alternative to LiDAR-based methods. However, directly leveraging implicit neural representations such as Neural Radiance Fields (NeRF) for detection faces fundamental challenges, including channel imbalance between RGB and density features and noisy density distributions that degrade localization accuracy. In this paper, we introduce two lightweight and complementary modules to enhance voxel feature quality for NeRF-based 3D detection. First, a geometry-aware fusion module processes appearance and density channels through separate modality-specific branches before recombining them via adaptive fusion, mitigating feature imbalance while amplifying geometric cues. Second, a contour-aware attention mechanism with a density-guided suppression loss reweights voxel features by emphasizing structural boundaries and penalizing background activations, yielding compact and morphology-consistent voxel fields. Extensive experiments on Hypersim, 3D-FRONT, and ScanNet demonstrate consistent improvements over the NeRF-RPN baseline in both recall and precision. Our modules add only 84 parameters and 0.328 GFLOPs, making them readily deployable within existing NeRF-based detection pipelines.
Results validate the effectiveness of the proposed novel 3D object detection and tracking framework, termed ECF3DMOT, in advancing 3D object detection and tracking for autonomous driving.
Xiaojuan Peng, Fei Teng, Tiankai Chen et al.· International Journal of Mac...· 0 citations
Multi-modal 3D object detection is an important task in autonomous driving systems, where cameras and LiDAR provide complementary semantic and geometric information. Most existing BEV fusion methods are designed based on the Cartesian representation space, which does not fully match the sensing geometry of camera and L...
Feng Gao, Jiaxin Chen, Niuniu Wang· Italian National Conference...· 0 citations
This work proposes a cascade optimization framework that systematically enhances feature representation and refines multimodal fusion, and introduces the Multi-Scale Contextual Fusion Module (MSCF) to reduce alignment bias.
Multimodal 3D object detection is fundamental to robust perception in autonomous driving because it integrates complementary information from LiDAR and camera sensors. However, existing methods often fail to maintain robustness under out-of-distribution (OOD) corruptions caused by sensor noise, adverse weather, and env...
Zi-Ying Song, Lin Liu, Hong-Yu Pan et al.· 0 citations
High-precision perception is fundamental to safe autonomous driving, and BEV-based 3D object detection via lidar-camera fusion plays a crucial role in improving detection accuracy and robustness. To address insufficient feature representation, spatial misalignment, and the limitations of static fusion strategies, this...
Jie Hu, Xinghao Cheng, Shuaidi He et al.· International Conference on...· 0 citations