Radar-Audio UAV Classification Under Asymmetric Degradation Using Training-Inference Decoupled Gating
Multimodal radar–audio sensing combines the spatial and reflectivity cues provided by millimeter-wave radar with the class-discriminative spectral patterns of acoustic signals for unmanned aerial vehicle classification. Under asymmetric degradation, however, a corrupted acoustic branch may become detrimental because continuously weighted acoustic features can remain in the fused representation. To investigate this failure mode, a controlled benchmark is constructed by applying Gaussian perturbations to stored Log-Mel representations while keeping the paired radar inputs unchanged. RAD-Gating is proposed as a reliability-aware fusion framework that decouples classifier-backbone training, reliability-gate calibration, and threshold-based inference. After the classifier backbone has been trained, its parameters are frozen and a lightweight reliability estimator is calibrated using polarized clean and degraded samples. During inference, the predicted acoustic reliability score is converted into a binary gate; unreliable acoustic features are zero-masked, and activation-scale compensation is applied to the retained radar representation. On the fixed 80/20 MMAUD split, the accuracy at −10 dB controlled feature-level degradation is increased from 29.23% with Softmax fusion to 56.15% with RAD-Gating, while the dedicated radar-only classifier achieves 61.54%. Fixed-checkpoint analyses are further conducted to examine reliability-score separation, threshold selection, feature-scale control, and soft versus hard inference. Within this controlled synthetic feature-level benchmark, the results indicate that reliability-aware hard suppression can mitigate multimodal negative transfer. These findings should not be interpreted as evidence of robustness to waveform-level or real-world acoustic disturbances.