Toward Automated and Personalized Employee Performance Evaluation: A Systematic Literature Review of NLG and LLM-Based Systems
Abstract
Organizations increasingly rely on heterogeneous digital performance data, including KPIs, activity logs, competency ratings, and 360-degree feedback, to support employee appraisal and development. However, transforming such evidence into interpretable, personalized, and governable narrative feedback remains an open challenge. Natural Language Generation (NLG) and large language models (LLMs) offer a promising basis for this transformation, yet no prior review has systematically examined their applicability to automated and personalized employee evaluation systems. This study presents a systematic literature review conducted using PRISMA-informed procedures. Searches across IEEE Xplore, ACM Digital Library, ScienceDirect, Scopus, and Web of Science identified studies published between January 2015 and September 2025. After screening and eligibility assessment, 98 papers were retained and analyzed across five dimensions: task families, methodologies, data sources, evaluation practices, and deployment challenges. As employee evaluation is an emerging application area, the retained studies are drawn primarily from adjacent domains, including education, healthcare, and coaching, and the review therefore synthesizes transferable components rather than direct evidence from workplace settings. The review identifies four task families relevant to employee evaluation: automated feedback generation, evaluative judgment, report and template generation, and personalized developmental guidance. Prompt-based LLM approaches dominate the methodological landscape. Personalization remains relatively shallow and is rarely validated against downstream outcomes. Governance concerns, particularly those related to privacy, fairness, and deployment-time explainability, are inconsistently addressed, while the evaluation methodology remains the main bottleneck in the field. In general, the necessary building blocks exist across adjacent domains but remain fragmented, and their transfer to workplace evaluation will require domain-specific adaptation. Progress will require integrated architectures combining evidence grounding, role-sensitive personalization, standardized evaluation protocols, and institutional accountability.