Lightweight Transformer-Fourier Fusion Framework for Efficient Image Super-Resolution
Abstract
Image super-resolution (ISR) has emerged as a critical computer vision task aimed at reconstructing high-resolution visual information from low-resolution inputs. Although deep learning-based approaches have significantly improved reconstruction quality, many existing architectures suffer from high computational complexity, excessive parameter requirements, and limited efficiency in real-time deployment scenarios. This research presents a Lightweight Transformer-Fourier Fusion Framework for Efficient Image Super-Resolution, designed to integrate the long-range dependency modeling capability of transformers with the frequency-domain representation advantages of Fourier-based feature processing. The proposed framework is theoretically positioned around efficient feature extraction, adaptive attention learning, and frequency-aware reconstruction. Transformer-based modules enhance spatial relationship modeling, while Fourier convolution mechanisms improve the preservation of high-frequency image details with reduced computational overhead. The methodology combines lightweight residual feature refinement, transformer-driven contextual enhancement, and frequency-domain fusion to achieve an optimized balance between reconstruction accuracy and computational efficiency. The study analyzes the limitations of conventional convolutional, generative, and attention-based super-resolution approaches and establishes the importance of hybrid architectures for next-generation ISR systems. The proposed framework provides a scalable solution for applications requiring efficient image enhancement, including mobile imaging, medical visualization, remote sensing, and intelligent surveillance systems.