A speech enhancement framework that jointly models Mel-spectrogram and complex spectrogram in a dual-path architecture for improved ASR performance and higher-quality speech reconstruction, and demonstrates consistent improvements in speech fidelity, perceptual quality, and ASR performance.
Abstract
In this paper, we propose DualSpecSE, a speech enhancement framework that jointly models Mel-spectrogram and complex spectrogram in a dual-path architecture for improved ASR performance and higher-quality speech reconstruction. The Mel branch learns coarse-grained acoustic representations and produces enhanced Mel-spectrograms for direct ASR usage, while the complex branch refines fine-grained spectral details for high-fidelity waveform reconstruction. Built upon the cross-band and narrow-band blocks from CleanMel, DualSpecSE introduces an interaction module and a fusion module to enable effective information exchange between the two branches. The model simultaneously outputs enhanced Mel and complex spectrogram without requiring a pretrained vocoder. Experimental results demonstrate consistent improvements in speech fidelity, perceptual quality, and ASR performance. Codes and audio samples are available.
Evaluation results demonstrate that DSVN improves restoration quality under both single- and mixed-degradation scenarios, and shows clear advantages in non-intrusive quality assessment, spectral distance, and automatic speech recognition evaluation, indicating that the enhanced speech is not only cleaner but also more...
Jun-Kang Yang, Hiromitsu Nishizaki, C. Leow et al.· Journal of the Acoustical So...· 0 citations
In this paper, we propose an advanced speech enhancement model capable of effectively separating clean speech from noisy audio signals. The primary objective here is to improve speech intelligibility and quality in noisy environments while preserving critical speech components. We propose a GAN based novel residual lea...
Debabrata Gogoi, Sushanta Kabir Dutta· Engineering Research Express· 0 citations
Increasing the overall realism of synthetic acoustic training data improves the generalization of DeepFilterNet3 to unseen measured environments and shows that increasing the overall realism of synthetic acoustic training data improves the generalization of DeepFilterNet3 to unseen measured environments.
Alessia Milo, G. Götz, S. Guðjónsson et al.· 1 citation
This paper presents MSFEYNet, a novel dual-decoder U-Net architecture for single-channel speech enhancement that integrates convolutional and Transformer blocks to exploit local spectral information and long-range contextual dependencies jointly. The encoder extracts hierarchical multi-scale spectral representations, w...
Silpa Peethala, S. Vanambathina· International journal of rec...· 0 citations
Speech enhancement plays a fundamental role in improving the perceptual quality and intelligibility of speech signals degraded by noise. Conventional U-Net architectures effectively capture local spectral information but struggle to model long-range dependencies and may allow residual noise to propagate through ski...
Shaik AreefaBegam, Sunnydayal Vanambathina· Frontiers in Signal Processi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.