Skip to content

Author

Tianping Li

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Jul 2026

SFC-DETR: An Asymmetric Spatial-Frequency Intervention Paradigm for Dense Aerial Object Detection

Unmanned aerial vehicle (UAV) object detection struggles with weak and tiny targets due to extreme viewing angles and environmental clutter. While the Real-Time Detection Transformer (RT-DETR) offers end-to-end efficiency, its macroscopic pooling and global attention mechanisms intrinsically discard the local spatial coordinates of tiny objects, causing severe missed detections in dense scenarios. To overcome this, we propose SFC-DETR, an efficient architecture featuring spatial-frequency decoupling and late feature compensation. Specifically, a Spatial-Frequency Decoupling Module (SFDE) is embedded in deep layers to filter low-frequency background noise and purify high-frequency target textures via spatial differential approximation. Serving as a pre-attention purifier, it significantly mitigates background crosstalk. Concurrently, to counteract deep-network resolution decay, a Feature Resolution Entry (FRE) module extracts high-fidelity microscopic features from extremely shallow layers for direct late compensation at the decoder's front end, thoroughly bridging the spatial representation gap. Extensive experiments on the VisDrone2019 dataset demonstrate that SFC-DETR achieves substantial gains in small object precision (mAP_50) and overall accuracy (mAP{50-95}) while maintaining real-time inference speeds, establishing a robust and highly efficient paradigm for weak object recognition in dense UAV scenarios.

Ningxiang Sun, Jie Li, Bing Liu et al. · 0 citations
Conference Jul 2026

Traffic Sign Detection Combining Frequency-Domain Enhancement and Local Channel Attention

Traffic sign detection represents a critical visual perception task in intelligent transportation systems and autonomous driving technologies, where accurate detection directly impacts driving safety. However, existing methods still face two major challenges in practical deployment: difficulty in small object detection and insufficient adaptability to complex environments. To address these issues, this paper proposes a traffic sign detection method integrating frequency domain enhancement and local channel attention based on RT-DETR. First, the backbone network structure is optimized by incorporating Cross Stage Partial (CSP) connections, achieving local-global-frequency domain three-dimensional collaborative feature extraction while maintaining comparable parameter count and computational overhead. Second, we design the EVCGLU (Enhanced Vision Convolutional Gated Linear Unit) module, which sequentially stacks five core components along the residual connection pathway: 3×3 Depthwise Separable Convolution (DWConv), Hidden State Mixer-based State Space Duality (HSM-SSD), Layer Normalization (LayerNorm), 3×3 Depthwise Separable Convolution, and Convolutional Gated Linear Unit (CGLU). This architecture implements local feature-based channel attention, effectively enhancing model robustness. Experimental results on the TT100K dataset demonstrate that the proposed method achieves mAP@0.5 of 83.1% and mAP@0.5-0.95 of 64.7%, representing improvements of 0.2 and 0.7 percentage points over the baseline RT-DETR-R18, respectively. The parameter count is merely 14.54M with computational cost of 48.1G FLOPs, reducing by 27.0% and 15.8% compared to RT-DETR-R18. The inference speed reaches 96 FPS, satisfying the real-time requirements of onboard embedded devices.

Fang Niu, Jia-Jing Sun, Shuang-Qiang Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.