Skip to content

Author

Xiaokai Bai

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Oct 2026

KCTF-Net: Kinematic Compensation Temporal Fusion for 4D Radar 3D Object Detection

4D millimeter-wave (mmWave) radar enables all-weather 3D object detection with reliable Doppler sensing. However, its practical application is hindered by inherent sparsity and noise. While multi-frame accumulation densifies point clouds, it inevitably introduces motion-induced spatiotemporal misalignment and geometric distortion, degrading detection accuracy in dynamic environments. To address these challenges, we propose the Kinematic Compensation Temporal Fusion Network (KCTF-Net), a novel 3D detection framework. Specifically, the Dynamic Kinematic Compensation (DKC) module explicitly aligns dynamic points in physical space, rectifying the motion-induced “smearing” effect. Furthermore, the Temporal Pillar Flow Enhancement (TPFE) module captures latent inter-frame kinematics, adaptively suppressing noise and mitigating sparsity. Finally, the Motion-Guided Attention Fusion (MGAF) module synergistically integrates density-filtered geometric priors with temporal features to reconstruct precise object geometries. Extensive experiments on the View-of-Delft (VoD) and TJ4DRadSet benchmarks demonstrate that KCTF-Net achieves state-of-the-art performance. Notably, it yields a 3D mean average precision (mAP) of 57.73%.

Xing-Kai Jin, Guang-Xian Xu, Fei Ma et al. · 0 citations
Preprint Jul 2026

4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction in 4D Radar-Camera Full-Scene Perception

Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout. Recently, 4D millimeter-wave radar has emerged as a robust and affordable sensor, yet its sparse returns make radar-camera fusion necessary for comprehensive scene understanding. Existing radar-camera methods mainly optimize detection, while dual-task systems usually decode boxes and occupancy with limited interaction. To address this gap and advance radar-based multi-task learning, we propose \method, a 4D radar-camera framework for 360$^\circ$ full-scene perception, which models semantic occupancy as a persistent scene state rather than a terminal output. \method{} follows a cross-modal state reasoning paradigm, where the occupancy state is modeled and propagated through stages for coarse-to-fine feature aggregation. Specifically, State-guided BEV Enhancement (SBE) strengthens intra-frame BEV representation, while Doppler-guided Temporal Fusion (DTF) preserves state evidence over longer temporal horizons. Beyond the model, we further extend ManTruckScenes with satellite-map-based generated occupancy labels and pair it with OmniHD-Scenes in a unified cross-dataset detection-and-occupancy protocol. The resulting experiments cover accuracy, robustness, ablation, and efficiency under one radar-camera multi-task evaluation framework. Code and labels will be released upon acceptance.

Xiaokai Bai, Lianqing Zheng, Runwei Guan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.