Back to feed
Open access

Metric-Driven Reinforcement Learning for Object Hallucination Reduction in LVLM-Based Image Captioning

2026 · IEEE Access · Vol 14, pp. 125177-125195 · 0 citations · 95 references

Abstract

Large Vision-Language Models (LVLMs) have excelled in joint visual and language understanding, particularly in image captioning. However, LVLMs remain prone to object hallucination, describing non-existent entities in the given image. Previous remedies, including dataset-enhanced supervised training and reinforcement learning from human feedback, show promise but rely heavily on costly human involvement, limiting scalability. To address this, we propose a reinforcement learning framework to mitigate object hallucination in LVLM-based image captioning, driven entirely by natural language processing (NLP) metrics. While optimizing reinforcement learning solely on NLP metrics offers scalability, it also introduces pitfalls: reward hacking, sparse feedback, and costly training. Our framework sidesteps these by introducing an Object Consistency Score (OCS) capturing both correctness and coverage of objects, incorporating KL-regularization for denser token-level guidance, and designing a lightweight PPO variant that fine-tunes adapters atop a frozen LVLM backbone, achieving efficiency and stability. Extensive experiments on the InstructBLIP baseline demonstrate the effectiveness and generalizability of our framework, achieving a substantial hallucination reduction of up to 41% on the COCO dataset and 29% on the Visual Genome. Furthermore, the framework can enhance caption quality by incorporating complementary evaluation metrics such as BERTScore and METEOR. These results reveal a broader insight in multimodal learning: LVLMs can effectively self-correct hallucinations using metric feedback, motivating deeper exploration into scalable reinforcement frameworks that significantly reduce human effort.

Read PDF