Skip to content

Author

Haopeng Zhang

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Channel Attention-Based Multi-Domain Feature Alignment for Moving Vehicle Detection in Satellite Videos Toward Smart Urban Planning

Rapid global urbanization is increasing the need for accurate, large-scale traffic monitoring to support sustainable transportation and city governance. Satellite video remote sensing offers a unique way to continuously observe urban road networks over large areas. It provides high-resolution spatio-temporal data that is essential for traffic flow analysis, infrastructure assessment, and dynamic urban planning. Moving vehicle detection in satellite video sequences is a basic task that turns raw imagery into useful traffic-state information, supporting these applications. Despite the advantages of satellite video data, detecting moving vehicles in practice remains a tough problem. Objects are extremely small and lack clear appearance details, while low local contrast makes them hard to separate from complex backgrounds. Satellite platform motion also introduces background misalignment and intensity fluctuations, resulting in missed detections and false alarms that hurt monitoring reliability. Furthermore, current methods do not fully exploit temporal motion cues or transform-domain priors, creating a performance bottleneck that restricts their practical use. To solve these problems, this paper proposes a Channel-Attentive Spatio-Temporal-Frequency Alignment (CASTFA) framework to effectively use and combine multi-dimensional features for moving vehicle detection in satellite videos, with the goal of providing high-quality traffic monitoring data to help smart city planning. Specifically, a State Space-Guided Temporal Compression (SSGTC) module first collects information along the time dimension with linear computational complexity, greatly reducing overhead while keeping motion cues that are critical for traffic-state estimation. The compressed temporal features are then processed with a multi-scale Haar wavelet transform to get hierarchical time-frequency representations that capture subtle motion dynamics across different frequency bands. At the same time, a pre-trained backbone network extracts multi-scale spatial features. To allow these different domains to work together, a Cross-Domain Feature Alignment (CDFA) mechanism aligns and combines spatial and time-frequency features through channel-attentive operations. Experimental results on the publicly available satellite video moving vehicle detection dataset show that the proposed CASTFA method consistently outperforms existing approaches, with better precision, recall, and F1-scores across diverse urban scenarios. These results show that CASTFA can provide reliable moving vehicle detection performance under difficult real-world conditions, supporting accurate traffic-flow monitoring and providing valuable geospatial intelligence for smart urban planning, transportation management, and sustainable city development.

Ning Zhao, Xiao Wang, Xiaopeng Zhang et al. · 0 citations
Preprint Aug 2026

CRISP: Calibration-Aware Visual State Space Duality for Remote Sensing Semantic Segmentation

State space models, especially Visual State Space Duality (VSSD), have emerged as efficient linear-time alternatives to Transformers for dense visual tasks. However, we observe that VSSD compresses spatial context into a global aggregation that suppresses high-frequency responses, causing excessive boundary smoothing in remote sensing semantic segmentation. To address this, we propose CRISP, a calibration framework with two components. Its core, the Duality Calibration Operator (DCO), restores local contrast and boundary responses through residual injection and frequency calibration within the VSSD backbone, without altering its linear complexity. To retain the recovered detail, an Orthogonal Multi-Prototype (OMP) head assigns multiple orthogonally constrained prototypes per class to model large intra-class variance. Extensive experiments on Potsdam, Vaihingen, and LoveDA show that, with approximately 30M parameters, CRISP achieves consistent gains in mean F1 (mF) and mIoU while remaining competitive with state-of-the-art methods. Code is available at https://github.com/crazylifeha/CRISP.

Kangning Wang, Haopeng Zhang, Zhiguo Jiang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.