Scene Text Retrieval (STR) aims to search images containing a given textual query within large-scale image collections. However, existing approaches are fundamentally constrained in two ways: 1) they are evaluated on narrow benchmarks that focus primarily on natural scenes; and 2) they rely on either error-prone multi-stage recognition-then-matching pipelines or localization-assisted matching strategies. To address these limitations, we introduce MuST, the first comprehensive benchmark for multi-scene and bilingual STR tasks, covering a broad range of real-world scenarios with carefully curated Chinese and English textual queries. On top of this benchmark, we propose BPE-Ret, a novel Byte-Pair Encoding (BPE)-level retrieval framework built on a simple yet powerful principle: a word is considered present in an image if and only if all of its constituent subwords are present. Concretely, BPE-Ret decomposes textual queries into BPE subwords and directly aligns them with dense visual features within a unified embedding space, thereby eliminating the need for explicit text spotting and coarse-grained word-level matching. We further enhance fine-grained alignment through two key innovations: a weighted preference learning scheme that prioritizes challenging cases to sharpen discrimination on confusable word-image pairs, and a subword inclusive-OR matching strategy that enforces constituent subword verification to enable robust retrieval beyond word-level granularity. Extensive experiments show that BPE-Ret establishes new state-of-the-art performance on both existing public STR benchmarks and our newly proposed MuST dataset, demonstrating the effectiveness and robustness of subword-level retrieval for real-world, multilingual scene text understanding.
Tong-Kun Guan, Yu-Tong Cai, Haocheng Wang et al.· IEEE Transactions on Image P...· 0 citations
VaRS-Doc is proposed, a visual document retrieval framework that diversifies document representations by enabling the model to actively explore variant latent interpretations during document encoding, while preserving efficient late-interaction retrieval in which each query adaptively selects the best-fit representation.
Haocheng Wang, Tongkun Guan, Wei Shen et al.· 0 citations
The veracious semantic alignment in autoformalization is significant for formal mathematical reasoning. However, existing evaluations provide only opaque binary verdicts or scalar scores, offering no interpretable insight into where or why translations fail. This opacity severely limits both human understanding and automated system improvement. To bridge this gap, we introduce FormalRx, a comprehensive diagnostic evaluation framework that transforms autoformalization assessment from black-box judgments into actionable feedback. At its core is SCI Error Taxonomy, a hierarchical classification scheme decomposing autoformalization errors into 28 distinct categories with strict priority ordering. Building on this taxonomy, FormalRx provides four critical diagnostic capabilities: alignment verdicts, error categorization, error localization, and correction. We instantiate the framework with a diagnostic model FormalRx-8B, trained on 56,287 NL-FL pairs with fine-grained diagnostic annotations, and release FormalRx-Test as the first fine-grained diagnostic benchmark. FormalRx-8B achieves F1-scores of 0.88 (verdict) and 0.71 (categorization), along with accuracies of 0.75 (localization) and 0.73 (correction), substantially outperforming both general-purpose LLMs and specialized baselines. By connecting evaluation with actionable insights, FormalRx enables systematic diagnosis and improvement of autoformalization systems.
Haocheng Wang, Baiyu Huang, Yingjia Wan et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.