MuCRE-TextVQA: Mamba-enhanced uncertainty-aware counterfactual relational executor for text visual question answering
Abstract
Text-based visual question answering (TextVQA) jointly interprets image content, scene text, and natural-language questions from both fixed-vocabulary and OCR-derived answer spaces. This study focuses on spatial-relation cases, where cross-modal alignment and relation drift are especially severe. MuCRE-TextVQA combines USG, CFT, Hybrid Mamba2, and CRE, with CRE serving as the main relation-execution component. With the updated results, MuCRE improves spatial-subset accuracy from 0.3774 to 0.3936 while reaching 0.4449 overall accuracy; CFT alone remains slightly higher in overall accuracy (0.4491). The claimed advantage is therefore relation-sensitive reasoning rather than uniformly best overall performance.