GS-IQA is proposed, a framework that reformulates IQA as a progressive Where--What--How diagnosis, emulating the human perceptual process from an initial glance to closer scrutiny, and consistently surpasses state-of-the-art methods in distortion localization, recognition, and severity estimation.
Abstract
Multi-modal large language models (MLLMs) have demonstrated significant potential in image quality assessment (IQA) by bridging visual perception with descriptive evaluations. However, existing approaches mainly focus on holistic quality prediction, often functioning as black boxes that provide limited insight into where distortions occur and how they affect perceived quality, hindering fine-grained analysis of localized and heterogeneous degradations. We propose GS-IQA, a framework that reformulates IQA as a progressive Where--What--How diagnosis, emulating the human perceptual process from an initial glance to closer scrutiny. Since a severity judgment is meaningful only for a correctly localized and recognized region, we realize this progression through a two-stage reinforcement learning paradigm that respects such dependencies: the glance stage uses a perception-gated reward to establish where degradations lie and what they are, activating severity feedback only once both are correct, while the scrutiny stage introduces online reward-conditioned degradation generation to synthesize hard examples targeted at the model's perceptual bottlenecks, sharpening its discrimination of subtle severity variations. To enable systematic evaluation, we construct Diag-Bench, a region-level IQA benchmark of about 25K curated samples spanning 12 distortion types and five ordinal severity levels. Extensive experiments show that GS-IQA consistently surpasses state-of-the-art methods in distortion localization, recognition, and severity estimation, and that its diagnostic representations transfer effectively to conventional global quality prediction across diverse external benchmarks. Code and data will be released.
Generative models are rapidly expanding image quality assessment (IQA) beyond traditional fidelity factors to emerging dimensions such as physical plausibility and text-rendering correctness. However, existing IQA models rely on fixed definitions and heavy supervision, making them difficult to extend to open-ended perc...
Zhen-Chen Tang, Bo Peng, Zi-Chuan Wang et al.· 1 citation
VGI-bench is introduced, containing 27 tasks and 810 instances, organized by a two-level taxonomy of task domains and skill tags for fine-grained evaluation of visual reasoning capabilities of video generation models, and it is hoped VGI-bench will help stimulate the development of next-generation video generation mode...
Xuan He, Cong Wei, Yuhao Cheng et al.· 2 citations
AI-generated image quality assessment (AIGIQA) requires jointly reasoning about perceptual fidelity and prompt alignment, two quality dimensions that are often treated as independent in existing AIGIQA models. However, by re-examining human ratings, we uncover a previously overlooked phenomenon: the two dimensions are...
Fine-grained visual perception in MLLMs is commonly improved by raising the resolution, but the added visual tokens inflate vision-encoding and language-model prefilling costs. We show that the two operations underlying fine-grained perception, localizing the region of interest (RoI) and recognizing its content, have d...
Yuheng Shi, Xiao-Huan Pei, Min-Jing Dong et al.· 0 citations
EviViT, a lightweight attachment that learns where a pretrained vision transformer should acquire detail, is introduced, a lightweight attachment that learns where a pretrained vision transformer should acquire detail without refitting.
Yao-Xin Niu, Zhangquan Chen, Yang Zhang et al.· 0 citations
With the rapid advancement of text-to-image (T2I) generation, robust evaluation becomes critical yet challenging, as traditional metrics fail to capture fine-grained alignment and generative artifacts. While large multimodal models (LMMs) are increasingly adopted as evaluators, existing benchmarks typically study seman...
Yu Zhao, Jia-Rui Wang, Hui-Yu Duan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.