Skip to content

Author

Hongyu Jiang

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Reinforcement Learning Methods and Optimization Strategies in Preference Alignment Techniques for Large Language Models

The growing pre-training scale has improved large language models' performance in language generation and knowledge representation. However, their training objectives remain limited to fitting data distributions, thus making it difficult to guarantee that the output satisfies human intentions and safety constraints. Therefore, the utilization of preference information to optimize model behavior and align generated content with human expectations has become a key challenge in the post-training phase of large language models. This paper investigates the development path of large language model preference alignment techniques, focusing on the training mechanism of Reinforcement Learning from Human Feedback (RLHF) and its three-stage process, including supervised fine-tuning, reward model training, and PPO-based policy optimization. On this basis, it compares the characteristics of Direct Preference Optimization (DPO), Reinforcement Learning from AI Feedback (RLAIF), and other methods in terms of training stability, data dependence, computational cost, and generalization ability. It further analyzes the alignment tax, which reflects the trade-off between improving model safety and preserving general capabilities during preference alignment. The results show that RLHF still possesses a higher upper bound for alignment in complex interactive scenarios, but its multi-stage training process brings high data and computational costs. In contrast, offline preference optimization methods such as DPO reduce training complexity and improve resource efficiency, but their performance remains constrained by preference data distribution and limited policy exploration capabilities.

Hongyu Jiang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.