Low-altitude unmanned-aerial-vehicle imagery exposes object detectors to a distinctive combination of tiny object footprints, dense instance layouts, abrupt scale variation, and weak texture under complex urban backgrounds. Existing detectors usually address these factors by adding larger backbones, denser feature pyramids, or heavier attention, but such independent additions often amplify background responses and dilute the fine localization cues needed by small targets. This paper proposes DFA-Det, a dynamic feature augmentation detector that treats low-altitude small-object detection as a coupled problem of context preservation, scale calibration, and content-aware refinement. The method first introduces a poly-kernel inception enhancement branch to preserve shallow structural details while expanding the effective receptive field. It then builds a Multi-Scale Interaction Encoder with Adaptive Feature Prior Learning and an Adaptive Feature Scaling Layer, where the latter contains a Bi-directional Channel Fusion Module that learns channel-wise evidence exchange between adjacent resolutions. Finally, a Hierarchical Refinement and Adaptive Fusion Module performs dynamic upsampling, semantic refinement, and adaptive fusion before the detection decoder. Experiments on the public VisDrone and CODrone benchmarks show that DFA-Det improves small-object precision, crowded-scene recall, and cross-scale robustness compared with representative two-stage, one-stage, transformer-based, and recent YOLO-family detectors. Extensive ablations, heatmaps, and qualitative comparisons indicate that the proposed modules cooperate as a coherent dynamic feature enhancement mechanism rather than isolated architectural attachments.
Robust object detection for autonomous driving requires perception models that remain reliable when visible imagery is degraded by darkness, glare, rain, fog, motion blur, or long-range small targets. Visible and thermal infrared cameras provide complementary evidence, yet many RGB–thermal detectors fuse modalities, mainly as aligned tensors, and may underuse relational structure in channel responses, spatial layouts, semantic scales, and modality-specific uncertainty. This paper presents TopoGraph-Fusion, a hierarchical graph-guided dual-modal object detector that formulates fusion as topology-aware reasoning rather than direct feature concatenation. The proposed framework builds a dual-stream backbone for RGB and thermal images, constructs channel-wise topology through a channel-topology graph aggregation module, derives relation-aware spatial and channel global attention from affinity graphs, and replaces fixed feature-pyramid communication with a Graph-Guided Feature-Pyramid Network. A topology-regularized detection objective further encourages stable cross-modal correspondence while suppressing noisy all-to-all connections. Experiments on M3FD, FLIR, RGBTDronePerson, and VEDAI512 cover road scenes, adverse illumination, drone–person perception, and aerial vehicle detection. Within this validation scope, the results and visual analyses indicate that topology-guided fusion improves small-object recall, cross-modal consistency, and robustness under modality imbalance.
Pu Yu, Yanshan Ma, Yuheng Li et al.· Symmetry· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.