The growing pre-training scale has improved large language models' performance in language generation and knowledge representation. However, their training objectives remain limited to fitting data distributions, thus making it difficult to guarantee that the output satisfies human intentions and safety constraints. Therefore, the utilization of preference information to optimize model behavior and align generated content with human expectations has become a key challenge in the post-training phase of large language models. This paper investigates the development path of large language model preference alignment techniques, focusing on the training mechanism of Reinforcement Learning from Human Feedback (RLHF) and its three-stage process, including supervised fine-tuning, reward model training, and PPO-based policy optimization. On this basis, it compares the characteristics of Direct Preference Optimization (DPO), Reinforcement Learning from AI Feedback (RLAIF), and other methods in terms of training stability, data dependence, computational cost, and generalization ability. It further analyzes the alignment tax, which reflects the trade-off between improving model safety and preserving general capabilities during preference alignment. The results show that RLHF still possesses a higher upper bound for alignment in complex interactive scenarios, but its multi-stage training process brings high data and computational costs. In contrast, offline preference optimization methods such as DPO reduce training complexity and improve resource efficiency, but their performance remains constrained by preference data distribution and limited policy exploration capabilities.
The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.
Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al.· Information and Software Tec...· 394 citations· ⚡54
The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.
Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al.· IEEE Transactions on Softwar...· 178 citations· ⚡14
This state-of-practice investigation was performed using a literature review followed by a multiple-case study approach and presents how inconsistency between managerial strategies and execution can lead to failure by means of a behavioral framework.
Carmine Giardino, Xiaofeng Wang, P. Abrahamsson· International Conference on...· 175 citations· ⚡19
Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.
M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al.· e-Informatica Software Engin...· 157 citations· ⚡17
This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.
Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al.· Empirical Software Engineeri...· 127 citations· ⚡15
The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.
Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al.· Journal of Systems and Softw...· 111 citations· ⚡8
What does it take to trust AI-driven HVAC optimization? Our AI Model Factory combines agents, machine learning, reinforcement learning and deterministic checks in a governed workflow designed for messy, real-world building data. The post We built an AI factory for HVAC control appeared first on GPT-Lab.
MIT News · Artificial Intelligence· news.mit.eduAug 18, 2026
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.
Microsoft Research Blog· microsoft.comJul 30, 2026
LLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment. The post EvoLib: Turning experience into evolving knowledge appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduMay 20, 2026