A Three-Dimensional Novel View Synthesis Method Based on Vision Transformer and Dual Residual Branches
Abstract
To address the problems of insufficient global structural consistency and local texture blurring in novel view synthesis under single-view conditions, this paper proposes a three-dimensional novel view synthesis method based on the fusion of a Vision Transformer and dual residual branches. The proposed method employs a Vision Transformer (ViT) to extract global features and capture long-range dependencies through a self-attention mechanism. Meanwhile, two complementary local branches are constructed. The RESFB module is designed based on Fast Fourier Convolution to fuse spatial-domain and frequency-domain information, while the RESTiedSE module introduces a TiedSE attention mechanism into the Res2Net framework to adaptively enhance key channel responses. The global and local features are fused in a multi-scale manner and combined with the NeRF volume rendering paradigm to generate novel views. Experiments conducted on the SRN-Chair and SRN-Car subsets of the ShapeNet dataset demonstrate the effectiveness of the proposed method. The results show that the proposed model achieves state-of-the-art SSIM and competitive PSNR, and effectively improves the clarity and structural fidelity of single-view novel view synthesis.