A cross-regime diagnostic framework for evaluating symbol consistency in MLLMs is proposed and its multi-stage failure mechanisms are revealed, offering a reusable diagnostic perspective for reliability assessment and symbol-capability improvement in MLLMs.
Abstract
Multimodal large language models (MLLMs) have achieved strong performance in general visual question answering, yet their visual faithfulness in handling discrete symbolic information remains insufficiently understood. Symbols such as numbers, time expressions, identifiers, license plates, and alphanumeric strings impose low semantic redundancy and strict character-level constraints, making overall VQA accuracy or isolated OCR-style evaluation inadequate for diagnosing model reliability. To address this gap, this paper proposes a cross-regime diagnostic framework for evaluating symbol consistency in MLLMs. Under a unified protocol, we evaluate five representative models, including BLIP-2, InstructBLIP, LLaVA, InternVL, and Qwen, across VQA, TextVQA, a self-constructed Symbol Subset, Regime 2a with uncontrolled generative symbol rendering, and Regime 2b with controlled clear-symbol grounding. We further introduce a 2a $\rightarrow 2$ b paired recovery analysis to distinguish rendering-sensitive errors caused by upstream symbol degradation from persistent errors that remain under clear visual evidence. Regime 2a is interpreted as an uncontrolled stress probe rather than a clean OCR benchmark, and its accuracy reflects both upstream rendering quality and downstream model reading. Experimental results show that general VQA accuracy can substantially overestimate symbol-centered reliability, especially for weaker models. Symbol consistency failures are not merely OCR recognition errors, but arise from the interaction of target-region binding, character-faithful transcription, answer completeness, and language-prior normalization. Although clear-symbol conditions improve stronger models, persistent failures remain in character-level grounding, target binding, and task following. This study separates symbol consistency from general VQA evaluation and reveals its multi-stage failure mechanisms, offering a reusable diagnostic perspective for reliability assessment and symbol-capability improvement in MLLMs.
Multimodal Large Language Models (MLLMs) can assign similar confidence to answers that fail for different reasons. We propose HalluPrism, a behavioral diagnostic that re-runs an answer after visual degradation, blank-image replacement, and grounding or relation checks. These targeted probes yield a signature over visual-perturbation sensitivity (V ), image-removal confidence retention (L), and grounding/relation-probe instability (A). Across 58K+ examples from four benchmarks and four MLLMs, image-removal confidence retention is most prevalent, while grounding/relation-probe instability better separates failure families. Only 18 of 48 source-target checks are diagonally aligned, so the coordinates should be interpreted jointly rather than as independent causal sources. With the dataset fixed, the joint signature improves failure-family AUROC from 0.634 to 0.769 on HallusionBench and from 0.707 to 0.817 on VizWiz, with smaller gains on POPE and VSR. In pooled XGBoost analysis, AUROC rises from 0.78 with scalar confidence to 0.95 with (V, L, A) and 0.97 when confidence is added. The same signature does not automatically improve correctness ranking. The three tested direct scalarizations can harm it. These results separate failure diagnosis from abstention scoring: multimodal uncertainty should characterize failure structure before it is used to decide whether to abstain or correct.
Aman Prakash, Sourish Dasgupta, Tanmoy Chakraborty· 0 citations
This work presents MMOOC, a large-scale benchmark for evaluating refusal and robust answering abilities of MLLMs, and introduces an LLM-as-a-Judge metric to assess the correctness of model reasoning.
LLM-based evaluators of natural language generation (NLG) quality are widely deployed as scoring tools and as automated training signals, yet the internal procedure by which they assign a rating remains poorly understood. We investigate this procedure mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline that produces paired clean and corrupt summaries with controlled error intensity and explicit token-level modification maps, and a four-experiment battery of causal tracing, logit-lens vocabulary projection, and attention-head knockout applied to Themis (Llama-3-8B) and Prometheus (Mistral-7B). Both evaluators implement a structured, coherent evaluation pipeline operating in two stages: below layer 15, attention performs local error comparison and routes the result to the final input position; above it, the MLP cascade integrates the signal and writes the rating, with the decision crystallizing in the residual stream at a sharp late layer (L = 26 on Themis, L = 25 on Prometheus). Furthermore, a base-model control at the same scale (Llama-3-8B) reproduces the routing architecture and crystallization but not the stage separation, isolating the two mechanisms that fine-tuning specifically installs, suppression of below-L15 MLP contribution at the last position and a two-layer advance of the crystallization depth, indicating that fine-tuning sculpts an existing substrate rather than building the pipeline from scratch. We release the source code and data at https://github.com/himil-v/judge-mech
BEAR-Bench (Bilingual Enterprise and Academic Reasoning), a self-contained, complex English-and-Russian benchmark comprising 1000 human-annotated questions based on text-rich business and scientific documents, is introduced, and existing hallucination detection methods are compared.
L. Chubarova, A. Kuleshova, D. P. Volkov et al.· 0 citations
The results show that LVLM predictions are sensitive to lexical, OCR-derived and stylistic cues, with injected surface signals causing substantial changes in model predictions despite unchanged underlying image-text relationships.
Zhanna Mukhametsharip, Vera Demberg, Varsha Suresh Saarland University et al.· 0 citations
Optical Character Recognition (OCR) is a key component in the digitization of historical archives. Recently, Vision-Language Models (VLMs) have emerged as strong alternatives to traditional OCR systems, achieving state-of-the-art performance on standard benchmarks. However, their suitability for archival transcription remains insufficiently understood. In this work, we benchmark traditional OCR systems and VLM-based approaches on the Berrutti dataset, a challenging collection of Uruguayan dictatorship-era documents derived from microfilm scans. While VLMs consistently outperform traditional methods in terms of Character Error Rate (CER) and Word Error Rate (WER), we show that these improvements hide a more complex picture. Through a detailed qualitative analysis, we uncover systematic failure modes that are invisible to standard metrics, including orthographic normalization, spurious content generation, and semantic substitutions that preserve fluency while altering meaning. Errors affecting named entities are particularly critical, as they can introduce substantial semantic distortions with minimal impact on CER and WER. These findings reveal a critical gap between quantitative OCR performance and transcription fidelity in real-world archival settings, and highlight the need for evaluation frameworks that go beyond character-level accuracy to capture the semantic reliability of generated transcriptions.
Marina Gardella, Camilo Mariño, Diego Belzarena et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.