Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to the wrong same-class instance. Generic visual-question-answering accuracy marks the response as wrong, while object-hallucination metrics may regard both the object and attribute as image-supported; neith...
Multimodal large language models generate natural-language responses from visual inputs, yet may mention objects absent from an image. In medication assistance, accessible perception, and environmental decision-making, such hallucinations can create real-world safety risks. We propose Semantic-Spatial Agreement Verific...
Zi-Heng Ren, Qian Gao, Jun Fan et al.· 0 citations
This paper proposes Dual-Stream Cross-Anchor Correction (DSCC), which injects object-level visual anchors into the language model itself during fine-tuning and reaches the long-caption, low-hallucination region.
LingKai Bu, Qian Gao, Jun Fan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.