Sparse Voxel Meets View: Bridging Local Geometry and Global Semantics via Linear Attention in Place Recognition
Abstract
Achieving accurate LiDAR-based place recognition is a crucial step towards reliable autonomous navigation, as it enables robust loop closure and global localization without GPS. However, existing 3D point cloud descriptors often give up fine local geometry for broad semantic context. We propose SBC-Net, a dual-branch architecture that integrates sparse voxel features with Bird’s-Eye View (BEV) representations via Sparse BEV Convolution (SBC), with the aim of producing highly discriminative and robust descriptors. A cross-attention fusion module injects BEV-derived global layout cues into the 3D sparse convolution branch early, enriching local geometry with complementary top-down context. In order to achieve efficient long-range interactions over large point clouds, we introduce Group Linear Attention with Reweighting (GLARE) to approximate transformer self-attention with linear complexity. The resulting global descriptor captures both local structural details and global semantics while remaining computationally scalable. Extensive experiments on the Oxford RobotCar benchmark, MulRan dataset, and the RoboLoc dataset demonstrate that SBC-Net consistently outperforms recent point cloud and BEV-based methods. These results confirm the robustness of SBC-Net and highlight the benefits of its multi-view feature fusion and efficient attention mechanisms for reliable place recognition.