Sep 2026· IEEE Transactions on Image Processing· Vol 35, pp. 10272-10285· 0 citations· 59 references
Medicine
Abstract
Recent advances in single image super-resolution (SISR) have leveraged convolutional neural networks (CNNs) and vision transformers to model pixel-level statistics, often relying on increasingly complex architectures to capture spatial correlations. However, these approaches generally overlook a fundamental distinction between human and machine perception, that is, the human visual system prioritizes shape-centric and semantically guided interpretation over raw pixel fidelity. To bridge this gap, we propose the View Large to Measure Shape Network (VLMSNet), a novel SISR framework inspired by the shape-centric strategy of the human visual system. Specifically, VLMSNet first employs a Group Mask Generator (GMG) to derive shape-aware guidance from large-receptive-field features, providing structurally informed cues for reconstruction. An Omni Attention Block (OAB) is further introduced to jointly model spatial and channel dependencies over broad contexts, so as to better represent complex structures and textures. In addition, VLMSNet adopts a dual-branch architecture with Foreground Feature Extraction (FFE) and Background Feature Extraction (BFE), explicitly separating salient object structures from less informative regions for more targeted and efficient feature learning. Extensive experiments both on standard SISR benchmarks and real-world scenarios demonstrate that VLMSNet achieves superior performance over existing state-of-the-art methods in most cases, confirming its robustness and practical applicability.
Texture recognition remains challenging for modern vision models because discriminative evidence is often carried by higher-order spatial statistics rather than by object shape alone. While Vision Transformers provide strong long-range modeling capacity, their standard object-centric representations do not explicitly e...
G-SalAlignMamba is proposed, a geometry-aware framework tailored for dual-modal SOD that introduces Geometry-Aware Encoding with explicit alignment to correct spatial shifts, Semantics-Informed Refinement to prevent signal dilution by prioritizing foregrounds, and Structure-Preserving Decoding that integrates explicit...
Hai-Xiao Gao, Yi-Min Zheng, Meng-Ke Song et al.· Proceedings of the Thirty-Fi...· 0 citations
VIVAS is proposed, a framework built upon the unified token space paradigm, which introduces a dense-structural-semantic vision tokenizer, which expands the textual vocabulary into a unified vision-language vocabulary by incorporating a visual vocabulary.
Zhe-Han Kan, Yu-Bo Zhu, Xing-Hua Jiang et al.· 0 citations
Vision-language models (VLMs) have emerged as powerful candidates for universal vision backbones, with representative architectures including autoregressive (AR) models and diffusion transformers (DiTs). Yet, adapting them efficiently for all-in-one low-level image restoration remains a challenge. Crucially, the field...
Shao-Jun Xia, Hui-Xin Zhang, Zhen Lei et al.· 0 citations
This work introduces SpatialBlock-15k, a synthetic dataset of 15,000 block-stacking problems covering 3D-to-2D projection, viewpoint transformation, and structural combination and proposes a novel paradigm inspired by human cognitive development: learning foundational spatial skills through structured block-manipulatio...
Soohyun Ryu, Sohee Kim, Eunho Yang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.