This work proposes RosePO, a framework to refine LLM-based recommendation through pairwise preference optimization with personalized smoothing, and incorporates a personalized smoothing factor predicted by a user oracle into the optimization objective.
Jiayi Liao, Xiangnan He, Ruobing Xie et al.· ACM Transactions on Informat...· 0 citations
This work proposes ARMOR (Anchor Rollout and Mixed Optimization for RL), a framework that shifts the paradigm from passive penalty to active sample stabilization, enabling sustained performance improvements over extended training horizons.
Kexin Huang, Junkang Wu, Jinda Lu et al.· arXiv.org· 0 citations
Perception-Enhanced Alignment DPO (PEA-DPO), a framework for multimodal LLMs alignment, which explicitly leverages visual preference signals to overcome visual insensitivity is proposed, which demonstrates that PEA-DPO enhances sensitivity to visual context while preserving the language modeling capacity of the base model.
Jiawei Feng, Jiancan Wu, Xingyu Zhu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.