DVFusion: Dual-View Attention and Geometry-Guided Fusion for 4D Radar-Camera 3D Detection
Abstract
4D radar-camera fusion has attracted increasing attention for reliable 3D perception under adverse weather conditions. However, existing methods either lose valuable radar information by using sparse point clouds, or suffer from high computation when processing raw tensors directly. To address these challenges, this paper proposes DVFusion, a progressive dual-view fusion framework that fully exploits the complementary geometric information from elevation-azimuth (EA) and range-azimuth (RA) representations of raw 4D radar tensors. DVFusion employs a geometry-aware progressive fusion strategy that first integrates image and EA-view radar features with focused linear attention, and subsequently fuses RA-view features in polar BEV space via geometry-guided fusion. In addition, a dual-dimensional adaptive residual mechanism is proposed to aggregate information across attention layers and adaptively balance feature contributions during fusion, alleviating feature dilution and enhancing representation capability. Experiments on the K-Radar dataset v1.0 show that DVFusion achieves 59.5% 3D mAP, outperforming existing methods under the same evaluation protocol. Additional experiments on K-Radar v2.0 confirm its ability to trade off detection accuracy against computational efficiency. This work suggests its promise for efficient and reliable all-weather 3D perception in autonomous driving.