This work constructs TextMuSS-10M, a large-scale synthetic scene text dataset spanning 10 scripts and 229 languages and proposes ScriptMoE, a script-aware Mixture-of-Experts (MoE) architecture that achieves the highest accuracy and is simpler than per-language experts, lighter than VLMs, and more accurate than both.
Xing-Song Ye, Yong-Kun Du, Jia-Xin Zhang et al.· 1 citation
DocIntent, a training-free Answerability-Guided Agentic Restoration framework, which first assesses question answerability, then identifies task-relevant degradations and selectively invokes restoration tools, and which consistently improves the average score and consistency of different open- and closed-source MLLMs.
Zi-Han Huang, Shi-Hang Wu, Jun-Le Liu et al.· 0 citations
This work proposes Single-Patch Text Spotting (SPaTS), a vision-centric framework that routes each text instance through a single anchor visual token and then recovers geometry via full-image refinement and introduces Single-Patch Selective Optimization (SPaSO), a reinforcement learning framework that optimizes discret...