Skip to content
Preprint

Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning

Aug 2026 · 0 citations · 46 references
Computer Science

TL;DR

An extended faithfulness analysis shows that the refined profiles remain largely grounded in the source preferences while preserving task-relevant personalization signals, suggesting that profile-side adaptation serves as a practical complement to universal memory construction for lifelong personalized agents.

Abstract

Natural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often contain information irrelevant to a particular downstream task. Directly supplying the full preference summary therefore wastes context capacity and introduces cross-task distraction, while manually designing task-specific preference views is difficult to scale. In this work, we study \emph{task-specific preference adaptation}: given a universal user preference summary and a downstream task, derive a task-conditioned representation that preserves sufficient decision-relevant evidence while removing redundant context. To this end, we propose \textsc{AlignXada}, a training-free meta-learning framework that induces reusable textual refinement policies for adapting universal preference summaries to task-specific ones. The refinement policy is iteratively optimized by a meta learner through verbal reinforcement learning. Across 13 tasks and three downstream models (39 task--model cells), \textsc{AlignXada} achieves an average gain of 3.82 points, improving 33 cells while retaining only 22.8\% of the original profile tokens and outperforming RAG in 36 cells. An extended faithfulness analysis further shows that the refined profiles remain largely grounded in the source preferences while preserving task-relevant personalization signals, suggesting that profile-side adaptation serves as a practical complement to universal memory construction for lifelong personalized agents.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Using Context Is Not Enough: Test-Time Training for Personalized Reward Modeling

Preference-Aligned Test-Time Training (P-TTT) is proposed, which explicitly encodes preference relations conveyed by contextual pairs into user-specific fast weights for personalized reward prediction and outperforms state-of-the-art methods by a large margin.

Bo-Hao Wang, Xiao-Yan Zhao, Yang Zhang et al. · 0 citations
Sep 2026

Instruction-Tuned LLMs via Bayesian Mixture Model Preference Priors

Large language models (LLMs) deployed as interactive recommendation agents must adapt to user preferences across dialogue rounds without access to ground-truth reward functions. Existing Bayesian-prompting approaches assume a uniform prior over user types, discarding population structure and weakening cold-start perfor...

Ornela Bregu, Nizar Bouguila · 0 citations
Preprint Aug 2026

Cautious Context Steering for Language Model Personalization

Cautious Context Steering (CCS), which adds a lightweight adapter to a frozen backbone LM to decide at each token whether and how strongly user context should affect generation, demonstrating robust generalization to new users and domains.

Gihoon Kim, Jeyoung Lee, S. Woo et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Less Is Personal: Learning Minimal Sufficient User Profiles for Personalized Language Models

Retrieval-augmented personalization enables large language models to produce more accurate and preference-aligned outputs using relevant records retrieved from user histories. Personalized language models typically prepend a fixed number of retrieved user records, even when additional history is redundant, harmful, or...

Ming-Hang Liu, Qiang Qiu, Yuan-Zhuo Wang et al. · 0 citations
#artificial intelligence Review Aug 2026

A Survey on Rubric-Guided Reinforcement Learning for Language Models

A Bayesian framework that defines constitutions as prior distributions over evaluation criteria and rubrics as conditional instantiations is introduced, and a taxonomy of rubric-guided RL along the prior-posterior axis is presented, covering constitutional AI, instance-specific rubrics, process-level supervision, self-...

Zifei Shan, Fang-Ning Shao · 1 citation
#artificial intelligence Preprint Aug 2026

Towards Reliable, Generalizable, and Specific In-Context Knowledge Editing via Multi-Objective Reinforcement Learning

Multi-Objective In-context Knowledge Editing (MO-IKE), a multi-objective RL algorithm that formulates prompt construction for in-context knowledge editing as a Constrained Markov Decision Process, enabling more balanced and globally coherent prompt construction.

Xu-Zhong Wang, Maiqi Jiang, Tejal Nair et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.