Perception-Enhanced Alignment DPO (PEA-DPO), a framework for multimodal LLMs alignment, which explicitly leverages visual preference signals to overcome visual insensitivity is proposed, which demonstrates that PEA-DPO enhances sensitivity to visual context while preserving the language modeling capacity of the base model.
Abstract
Direct Preference Optimization (DPO) has emerged as an effective approach for aligning large language models (LLMs) with human preferences. However, its adaptation to multimodal settings remains unexplored. Through representational analysis, we identify a key limitation in multimodal preference optimization, which we term visual insensitivity: models often fail to distinguish between images and those with critical visual context removed. Our theoretical analysis further uncovers two manifestations of this problem, namely Across-Image Insensitivity and Within-Image Insensitivity. To address these challenges, we propose Perception-Enhanced Alignment DPO (PEA-DPO), a framework for multimodal LLMs alignment, which explicitly leverages visual preference signals to overcome visual insensitivity. We further provide a theoretical analysis demonstrating that PEA-DPO provably mitigates both failure modes. Empirical results demonstrate that PEA-DPO enhances sensitivity to visual context while preserving the language modeling capacity of the base model. Evaluations across three hallucination benchmarks using MLLMs of varying scales show that PEA-DPO effectively mitigates visual insensitivity, achieves stronger multimodal alignment, and substantially reduces hallucinations.
Extensive experiments show that Groc-PO achieves improved performance in hallucination mitigation, faithful reasoning, and overall reliability, supporting the value of more explicit grounded supervision for trustworthy multimodal reasoning.
Zhi Zheng, Zheren Fu, Zhiyuan Yao et al.· arXiv.org· 1 citation
Context-Calibrated DPO (C$^2$-DPO), which directly maximizes CPG while preserving the original preference ordering, is proposed, which substantially reduces hallucination without compromising general reasoning.
Byungoh Ko, Jinyoung Park, Jongha Kim et al.· 0 citations
This work systematically characterize the implicit biases introduced by low-rank adaptation during alignment and establishes two theorems showing that low-rank alignment induces preferences for parameter subspaces with flat gradients and feature subspaces robust to perturbations, providing a principled explanation for the observed structure-preserving behavior.
Mingjia Shi, Shuo Wang, Xiaobo Wang et al.· arXiv.org· 0 citations
This empirical research establishes the essential groundwork for predictably scaling multimodal foundation models by modeling the influence of data composition on compute laws and allocation exponents and derive an efficiency frontier specifying precise configurations of model size, token count, and data mixture.
The core of SSVAL is Visual Anchor Prompt Injection (VAPI), which introduces prompts that absorb rich knowledge from external VFMs during training, enabling them to serve as stable visual anchors that mitigate representation deviation during inference.
Qian-Long Yang, Bowen Ye, Xianda Guo et al.· 0 citations
This paper formalizes preference privacy, a label-DP-style privacy notion for DPO that protects only the relative preference between candidate responses, assuming an adversary who already knows the prompt and responses, and designs PrivDPO, a DPO variant that enforces preference privacy while remaining compatible with large-scale LLM training.
Yangfan Jiang, Fei Wei, Ergute Bao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.