Skip to content

GlyphTSR: Text-Rich Scene Image Super-Resolution Beyond Glyph Priors

Aug 2026 · IEEE Transactions on Image Processing · Vol 35, pp. 9272-9284 · 0 citations · 53 references
Medicine

Abstract

Text-rich scene image super-resolution (TS-ISR) aims to recover high-quality images with legible text from degraded inputs, benefiting mobile photography and enhancing visual inputs for multimodal understanding. Existing methods rely on real-world image super-resolution or text image super-resolution, making it difficult to simultaneously preserve holistic scene fidelity and detailed character structure. To address the problem, this paper proposes a unified framework GlyphTSR that leverages glyph priors to enhance local glyph restoration while ensuring global semantic consistency. Specifically, the character aware extractor (CAE) with text-spotting guidance and the glyph region enhancer (GRE) with dual-tower encoder are designed to mine glyph features from low-quality inputs. CAE implicitly extracts latent glyph-related semantics to facilitate glyph restoration. GRE explicitly utilizes glyph and position cues to upgrade feature representation. The glyph-guided super-resolution with glyph guider and text-imbalance responsive loss (TIRL) is proposed to enhance local text restoration while performing global scene super-resolution. Furthermore, a new TS-ISR benchmark is established to jointly evaluate models’ ability to restore faithful scenes and legible text. Experiments on the proposed benchmark and public datasets demonstrate that the proposed GlyphTSR achieves competitive performance, establishing a baseline for the emerging TS-ISR. The benchmark is publicly available at https://github.com/qyx596/tsisr-benchmark

View source

Similar papers

Preprint Sep 2026

GlyphAnchor: Enhancing Visual Text Rendering via Position-Anchored Glyph Priors

Rendering accurate text remains difficult for image generation and editing models, especially when the target contains long, complex, and densely arranged text or rare characters. Existing approaches either improve native text rendering through stronger backbones and data-centric training without explicit glyph priors,...

Qiang Xiang, Shu'ang Sun, Binglei Li et al. · 1 citation
Preprint Aug 2026

MagnifiQ: Patch-aware Text Guided Progressive Upscaling for High-Resolution Image Restoration

MagnifiQ is introduced, an image restoration framework that progressively upscales and restores images across resolutions, e.g., from 1024x1024 to 4096x4096, and leverages a pre-trained text-to-image diffusion model such as SDXL and adapts it for more scalable high-resolution inference by replacing its original self-at...

M. Reddy, Yashesh Savani, Antoine Mercier et al. · 0 citations
Preprint Aug 2026

Coupled Continuous-Discrete Generation for Scene Text Image Super-Resolution

Scene text image super-resolution (STISR) aims to recover visually plausible appearance while preserving character semantics from degraded inputs. Existing STISR systems often rely on externally generated priors or separate image and text models, resulting in error propagation and costly multi-stage inference. We prese...

Axi Niu, Knag Zhang, Qing-Sen Yan et al. · 0 citations
Conference Open access Sep 2026

Auxiliary text-guided image restoration for image-text matching

Image-Text Matching (ITM) aims to establish deep semantic associations between visual content and textual descriptions. Existing methods usually have discrimination issues because of fine-grained semantic deviations, so it's hard to capture the complex correspondences between cross-modal entries. Only relying on alignm...

Kuang-Rong Hao · 0 citations
Aug 2026

STAFuse: Scene-Text Aggregation-Guided Composite Degradation-Robust Infrared and Visible Image Fusion

Infrared and visible image fusion aims to integrate complementary information from source images to generate high-quality fusion images that serve downstream tasks. However, the differentiated representation of image scene content, the unpredictability of degradation modes in source images, and the complexity of compos...

Ting Lv, Hong Jiang, Yu Liu · 0 citations
Preprint Aug 2026

TextRefine: Improving Textual Fidelity, Spatial Placement, and Glyph Rendering for Text Editing in Product Posters

Text editing in product posters entails inserting new text or replacing existing text while preserving product appearance, background content, and global composition. Despite recent progress in instruction-based image editing, general-purpose models remain unreliable in this setting: they often omit or incorrectly rend...

Hong-Lie Wang, Jia Sun, Zi-Jun Li et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.