Aug 2026· IEEE Transactions on Image Processing· Vol 35, pp. 9272-9284· 0 citations· 53 references
Medicine
Abstract
Text-rich scene image super-resolution (TS-ISR) aims to recover high-quality images with legible text from degraded inputs, benefiting mobile photography and enhancing visual inputs for multimodal understanding. Existing methods rely on real-world image super-resolution or text image super-resolution, making it difficult to simultaneously preserve holistic scene fidelity and detailed character structure. To address the problem, this paper proposes a unified framework GlyphTSR that leverages glyph priors to enhance local glyph restoration while ensuring global semantic consistency. Specifically, the character aware extractor (CAE) with text-spotting guidance and the glyph region enhancer (GRE) with dual-tower encoder are designed to mine glyph features from low-quality inputs. CAE implicitly extracts latent glyph-related semantics to facilitate glyph restoration. GRE explicitly utilizes glyph and position cues to upgrade feature representation. The glyph-guided super-resolution with glyph guider and text-imbalance responsive loss (TIRL) is proposed to enhance local text restoration while performing global scene super-resolution. Furthermore, a new TS-ISR benchmark is established to jointly evaluate models’ ability to restore faithful scenes and legible text. Experiments on the proposed benchmark and public datasets demonstrate that the proposed GlyphTSR achieves competitive performance, establishing a baseline for the emerging TS-ISR. The benchmark is publicly available at https://github.com/qyx596/tsisr-benchmark
Rendering accurate text remains difficult for image generation and editing models, especially when the target contains long, complex, and densely arranged text or rare characters. Existing approaches either improve native text rendering through stronger backbones and data-centric training without explicit glyph priors,...
Qiang Xiang, Shu'ang Sun, Binglei Li et al.· 1 citation
MagnifiQ is introduced, an image restoration framework that progressively upscales and restores images across resolutions, e.g., from 1024x1024 to 4096x4096, and leverages a pre-trained text-to-image diffusion model such as SDXL and adapts it for more scalable high-resolution inference by replacing its original self-at...
M. Reddy, Yashesh Savani, Antoine Mercier et al.· 0 citations
Scene text image super-resolution (STISR) aims to recover visually plausible appearance while preserving character semantics from degraded inputs. Existing STISR systems often rely on externally generated priors or separate image and text models, resulting in error propagation and costly multi-stage inference. We prese...
Axi Niu, Knag Zhang, Qing-Sen Yan et al.· 0 citations
Image-Text Matching (ITM) aims to establish deep semantic associations between visual content and textual descriptions. Existing methods usually have discrimination issues because of fine-grained semantic deviations, so it's hard to capture the complex correspondences between cross-modal entries. Only relying on alignm...
Kuang-Rong Hao· International Conference on...· 0 citations
Infrared and visible image fusion aims to integrate complementary information from source images to generate high-quality fusion images that serve downstream tasks. However, the differentiated representation of image scene content, the unpredictability of degradation modes in source images, and the complexity of compos...
Ting Lv, Hong Jiang, Yu Liu· IEEE Transactions on Image P...· 0 citations
Text editing in product posters entails inserting new text or replacing existing text while preserving product appearance, background content, and global composition. Despite recent progress in instruction-based image editing, general-purpose models remain unreliable in this setting: they often omit or incorrectly rend...
Hong-Lie Wang, Jia Sun, Zi-Jun Li et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.