Skip to content
Open access

TF-Mamba-DPHNet: a time-frequency state-space and dual-path interaction framework with hybrid cross-scale feature calibration for speech enhancement

Sep 2026 · Frontiers in Signal Processing · 0 citations · 55 references

Abstract

Speech enhancement plays a fundamental role in improving the perceptual quality and intelligibility of speech signals degraded by noise. Conventional U-Net architectures effectively capture local spectral information but struggle to model long-range dependencies and may allow residual noise to propagate through skip connections. Transformer-based networks can model global contextual dependencies and generate high-quality enhanced speech; however, their limited capacity to preserve fine-grained high-frequency spectral details can limit their suitability for real-time applications. To address these limitations, this paper proposes TF-Mamba-DPHNet, a novel encoder-decoder speech enhancement framework integrating Multi-Scale Feature Extraction (MSFE), Time-Frequency Mamba (TF-Mamba), Dual-Path Higher-Order Information Interaction with Time-Frequency Attention (DPH-TFA), and Hybrid Cross-Scale Feature Calibration (H-CS-FC) modules. The MSFE blocks in the encoder and decoder extract local patterns across multiple receptive fields, capturing fine-grained and global time-frequency cues. The TF-Mamba module models global time-frequency dependencies using selective state-space modeling for effective long-term sequence understanding. At the bottleneck, four stacked DPH-TFA blocks capture long-range dependencies in both the temporal and spectral domains. These global features are fused with hierarchical encoder outputs through H-CS-FC modules, which perform cross-scale-guided feature recalibration to suppress noise leakage in the skip pathways and improve decoder reliability. Evaluations conducted on the Common Voice and LibriSpeech datasets demonstrate that TF-Mamba-DPHNet effectively improves speech quality and robustness compared with earlier models. The results indicate that combining multi-scale local feature extraction, selective state-space modeling, dual-path time-frequency interaction, and cross-scale feature calibration enables effective modeling of complementary local and global information for robust speech enhancement.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.