RC-HyFusion: Radar–Camera Hybrid Fusion for 3-D Object Detection in Autonomous Driving
Abstract
In autonomous driving, achieving accurate and robust 3-D perception through the fusion of multiple sensor modalities is a critical requirement. While camera-based methods operating in the bird’s-eye view (BEV) have shown significant progress, they often suffer from performance degradation under adverse lighting and weather. Furthermore, these vision-based systems struggle with the reliable detection of distant or obscure objects. In this article, we propose radar–camera hybrid fusion (RC-HyFusion), a camera–radar fusion framework that achieves multimodal multistage fusion to generate a better feature representation. First, to address the depth ambiguity in images, sparse but accurate radar points are expanded into pillars in the image dimension to form pseudoimages. Next, multimodal features extracted from the radar pillars and images are sent to the radar–camera adaptive weighted fusion module, where interactive fusion is performed to generate fusion features. Then, fusion features are transformed to the BEV with a view transformation. Meanwhile, the radar-based adaptive sampling module aims to enrich the semantic information of radar points and realize multistage interaction by automatically sampling fusion features based on radar point features under the BEV. Finally, we utilize multimodal BEV features for 3-D detection. The experimental results demonstrate that RC-HyFusion effectively integrates radar and camera information and achieves competitive 3-D object detection performance on the nuScenes benchmark.