Skip to content
Open access

MSFEYNet: A Hybrid Convolution Transformer Network with Spectral Feature Extraction & MSFEB Modelling for Single-Channel Speech Enhancement

Sep 2026 · International journal of recent technology and engineering · 0 citations · 15 references

Abstract

This paper presents MSFEYNet, a novel dual-decoder U-Net architecture for single-channel speech enhancement that integrates convolutional and Transformer blocks to exploit local spectral information and long-range contextual dependencies jointly. The encoder extracts hierarchical multi-scale spectral representations, while two dedicated decoders simultaneously reconstruct the clean speech magnitude and phase spectra, enabling accurate speech recovery. To enhance feature representation, the proposed architecture incorporates three key components: the Triple Path Fused Attention (TPFA) module for efficient long-range axial spectral modelling, the Large Group Large Kernel Attention (LGLKA) module for capturing finegrained local axial structures, and an Enhanced Feed Forward Network (EFFN) with dynamic axial-aware feature reweighting to improve contextual feature learning. Furthermore, a Multi-Scale Feature Extraction Block (MSFEB) extracts both fine-grained local features and broad contextual information using convolutional filters with varying receptive fields, thereby enhancing multi-scale spectral-temporal feature representation. Extensive experiments on two datasets demonstrate that MSFEYNet achieves robust, well-balanced performance across a comprehensive set of objective evaluation metrics through the synergistic integration of the dual-decoder framework. The transformer-based attention mechanism and Multi-scale Feature Extraction Block (MSFEB) effectively capture local and global contextual dependencies, leading to significant improvements in speech intelligibility and perceptual quality. Experimental results further demonstrate that MSFEYNet consistently outperforms contemporary state-of-the-art speech enhancement methods, achieving substantial gains in STOI (Short-Time Objective Intelligibility) and PESQ (Perceptual Evaluation of Speech Quality) while maintaining excellent generalisation across diverse acoustic conditions and varying noise levels.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.