Jul 2026
Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs
Extensive experiments show that Groc-PO achieves improved performance in hallucination mitigation, faithful reasoning, and overall reliability, supporting the value of more explicit grounded supervision for trustworthy multimodal reasoning.
Zhi Zheng, Zheren Fu, Zhiyuan Yao et al.
· arXiv.org · 1 citation