Skip to main content

GlyphTSR: Text-Rich Scene Image Super-Resolution Beyond Glyph Priors

By
Na Jiang; Yuxuan Qiu; Jiawei Zhang; Xiangcheng Zhai; Dongqing Zou; Jimmy S. Ren; Yuan Xiong; Xiaochun Cao

Text-rich scene image super-resolution (TS-ISR) aims to recover high-quality images with legible text from degraded inputs, benefiting mobile photography and enhancing visual inputs for multimodal understanding. Existing methods rely on real-world image super-resolution or text image super-resolution, making it difficult to simultaneously preserve holistic scene fidelity and detailed character structure. To address the problem, this paper proposes a unified framework GlyphTSR that leverages glyph priors to enhance local glyph restoration while ensuring global semantic consistency. Specifically, the character aware extractor (CAE) with text-spotting guidance and the glyph region enhancer (GRE) with dual-tower encoder are designed to mine glyph features from low-quality inputs. CAE implicitly extracts latent glyph-related semantics to facilitate glyph restoration. GRE explicitly utilizes glyph and position cues to upgrade feature representation. The glyph-guided super-resolution with glyph guider and text-imbalance responsive loss (TIRL) is proposed to enhance local text restoration while performing global scene super-resolution. Furthermore, a new TS-ISR benchmark is established to jointly evaluate models’ ability to restore faithful scenes and legible text. Experiments on the proposed benchmark and public datasets demonstrate that the proposed GlyphTSR achieves competitive performance, establishing a baseline for the emerging TS-ISR. The benchmark is publicly available at https://github.com/qyx596/tsisr-benchmark

Read on IEEE Xplore