Skip to content
Conference

Graph-based Multi-Agent LLM Framework for OCR-based Visual Question Answering

Aug 2026 · International Conference on Multimedia Analysis and Pattern Recognition · pp. 724-729 · 0 citations · 44 references

Abstract

Recent advances in Large Language Models (LLMs) have improved reasoning in multimodal tasks such as Visual Question Answering (VQA). However, in OCR-centric scenarios such as signboard VQA, existing approaches remain vulnerable to hallucination, weak verification, and inconsistent reasoning when integrating visual and textual cues. In this paper, we propose a graph-based multi-agent LLM framework that introduces a structured intermediate reasoning layer between perception and language reasoning. OCR entities and visual objects are organized into a lightweight, layout-aware graph, and heterogeneous LLM agents (GPT and Gemini) exchange information exclusively through this shared structure rather than through free-form text, with a dedicated verification agent scoring and pruning hypotheses against the graph before answer generation. On ViSignVQA, a Vietnamese signboard VQA benchmark, our framework attains 55.80% F1 and 23.48% EM, improving over the strongest reported baseline by 4.04 F1 and 5.40 EM points. An ablation study shows that both modalities are required: removing OCR nodes or visual nodes reduces EM to 2.09% and 2.07%, respectively. On EVJVQA, a multilingual benchmark that is not OCR-centric, the framework transfers without collapsing, reaching the second-highest F1 (0.3618) among submitted systems, although its BLEU remains low — a gap we analyze as a property of LLM-generated answers rather than of reasoning quality. Our results highlight both the value and the current limits of structured reasoning and verification in multimodal LLM systems.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.