Results indicate that the proposed framework provides an effective solution with substantially reduced boundary discontinuities for advanced neural image compression systems based on hybrid CNN-Transformer architectures.
Abstract
We propose a high-capacity, end-to-end framework for large-scale image compression that addresses the trade-off between tiling scalability and perceptual quality, a challenge stemming from the patch-based processing required for high-resolution inputs, which often introduces disruptive stitching artifacts. To mitigate this issue, we present a unified framework built on three complementary components: 1) Accelerated Virtual-Tiling, which simulates boundary interactions during training to improve spatial consistency without incurring the memory cost of multi-patch encoding; 2) Seam-Targeted Distance-Masked Self-Attention, a latent bottleneck mechanism that enables information exchange across patch boundaries; and 3) Boundary-Aware Regularization, which enforces consistency at tile interfaces through an explicit loss formulation. By explicitly modeling cross-boundary dependencies, the proposed method effectively suppresses stitching artifacts while maintaining scalability to high-resolution inputs. Extensive experiments on the Kodak, JPEG AI, and CLIC 2025 datasets demonstrate competitive or superior rate-distortion performance, achieving high structural fidelity with MS-SSIM values of approximately 0.998 at high compression ratios. These results indicate that the proposed framework provides an effective solution with substantially reduced boundary discontinuities for advanced neural image compression systems based on hybrid CNN-Transformer architectures.
This work proposes a novel dual-branch architecture that integrates U-Net and Transformer networks: U-Net is leveraged for restoring fine-grained textures, while the Transformer effectively models long-range dependencies to ensure global semantic consistency.
Huaming Liu, Minglong Zhang, Xiuyou Wang et al.· Multimedia Systems· 0 citations
DPCA-Net establishes a superior Pareto front between restoration fidelity and efficiency by maintaining real-time inference speeds with a minimal footprint of 15.00 GFLOPs, offering a compelling practical solution for real-time all-in-one image restoration.
The results show that careful tokenizer--backbone--system co-design can deliver strong high-resolution generation and editing within an efficient 4B model family.
FIT employs a lightweight Degradation Encoder to predict a global degradation vector and a spatial degradation map from local degradation severity, which jointly condition the patch embedding and unembedding through adaptive deformation, and introduces a task-token dropout strategy that regularizes task conditioning during training.
Zihao He, Yunfeng Wu, Xinchao Wang et al.· 0 citations
TAMF-VTON is presented, a texture-aware, mask-free framework that enables high-fidelity image synthesis under practical unconstrained conditions and outperforms state-of-the-art methods in both quantitative metrics and perceptual quality.
Jie Wang, Qian He, Gaofeng He et al.· arXiv.org· 0 citations
An adaptive multi-scale decoding framework that effectively balances global context with fine-grained detail is proposed that exhibits superior robustness and generalization across diverse domains, effectively alleviating limitations of existing fusion-based approaches.