TF-Mamba-DPHNet: a time-frequency state-space and dual-path interaction framework with hybrid cross-scale feature calibration for speech enhancement
Abstract
Speech enhancement plays a fundamental role in improving the perceptual quality and intelligibility of speech signals degraded by noise. Conventional U-Net architectures effectively capture local spectral information but struggle to model long-range dependencies and may allow residual noise to propagate through skip connections. Transformer-based networks can model global contextual dependencies and generate high-quality enhanced speech; however, their limited capacity to preserve fine-grained high-frequency spectral details can limit their suitability for real-time applications. To address these limitations, this paper proposes TF-Mamba-DPHNet, a novel encoder-decoder speech enhancement framework integrating Multi-Scale Feature Extraction (MSFE), Time-Frequency Mamba (TF-Mamba), Dual-Path Higher-Order Information Interaction with Time-Frequency Attention (DPH-TFA), and Hybrid Cross-Scale Feature Calibration (H-CS-FC) modules. The MSFE blocks in the encoder and decoder extract local patterns across multiple receptive fields, capturing fine-grained and global time-frequency cues. The TF-Mamba module models global time-frequency dependencies using selective state-space modeling for effective long-term sequence understanding. At the bottleneck, four stacked DPH-TFA blocks capture long-range dependencies in both the temporal and spectral domains. These global features are fused with hierarchical encoder outputs through H-CS-FC modules, which perform cross-scale-guided feature recalibration to suppress noise leakage in the skip pathways and improve decoder reliability. Evaluations conducted on the Common Voice and LibriSpeech datasets demonstrate that TF-Mamba-DPHNet effectively improves speech quality and robustness compared with earlier models. The results indicate that combining multi-scale local feature extraction, selective state-space modeling, dual-path time-frequency interaction, and cross-scale feature calibration enables effective modeling of complementary local and global information for robust speech enhancement.